← blog
technical20 Sept 2026· 7 min read

Gradient descent, one step at a time

cover — live gradient descent

Almost every model I train in the MSc is fitted by some version of gradient descent. It's a short loop, and the learning rate changes how it behaves more than any other setting. Below it runs live, so you can try the learning rate yourself.

Picture the loss as a landscape. Every point is one choice of model parameters, and the height is how wrong the model is with them. Training means walking downhill. We can't see the whole landscape, only the slope where we're standing, so each step goes a little way in the steepest downhill direction.

for step in range(n_steps):
    grad = gradient(loss, θ)
    θ = θ - η * grad       # η = learning rate

Try it

The rings are contour lines of a stretched bowl, f(x, y) = ½(x² + 8y²), and the bright star at the centre is the minimum. The bowl is steep in one direction and shallow in the other, which is common in real models. Pick a learning rate and watch the path.

η =
step 0loss —running…
Fig. 1 — 40 steps from the same starting point. Choose a learning rate above.

What the four settings show

At 0.02 every step is safe but tiny, and 40 steps are not enough to arrive. At 0.1 it gets there cleanly. At 0.23 the steep direction overshoots: the path zig-zags across the valley and settles slowly. At 0.26 each overshoot is bigger than the last, and the loss blows up.

The limit is set by the steepest direction. For this bowl the curvature along y is 8, so any learning rate above 2 / 8 = 0.25 diverges. The shallow direction would happily take bigger steps, but it can't. That mismatch is why optimisers like momentum and Adam exist, and I'll cover them in the next post.

Shamsunnur Ibn Arefin is a data science student and former QA engineer, writing from Skövde, Sweden.

about →