Almost every model I train in the MSc is fitted by some version of gradient descent. It's a short loop, and the learning rate changes how it behaves more than any other setting. Below it runs live, so you can try the learning rate yourself.
Picture the loss as a landscape. Every point is one choice of model parameters, and the height is how wrong the model is with them. Training means walking downhill. We can't see the whole landscape, only the slope where we're standing, so each step goes a little way in the steepest downhill direction.
for step in range(n_steps):
grad = gradient(loss, θ)
θ = θ - η * grad # η = learning rate
Try it
The rings are contour lines of a stretched bowl, f(x, y) = ½(x² + 8y²), and the bright star at the centre is the minimum. The bowl is steep in one direction and shallow in the other, which is common in real models. Pick a learning rate and watch the path.
What the four settings show
At 0.02 every step is safe but tiny, and 40 steps are not enough to arrive. At 0.1 it gets there cleanly. At 0.23 the steep direction overshoots: the path zig-zags across the valley and settles slowly. At 0.26 each overshoot is bigger than the last, and the loss blows up.
The limit is set by the steepest direction. For this bowl the curvature along y is 8, so any learning rate above 2 / 8 = 0.25 diverges. The shallow direction would happily take bigger steps, but it can't. That mismatch is why optimisers like momentum and Adam exist, and I'll cover them in the next post.