06 · Neural nets · 5 min read · Interactive · updated
What is the learning rate and why does setting it too high break training?
In short
The learning rate scales the step along the negative gradient when updating weights. Too small slows training; too large causes oscillation and divergence.
What it is
The learning rate (written η or α) is the number by which we multiply the gradient of the loss before subtracting it from the weights: weight ← weight − η × gradient. The gradient points in the direction of the steepest increase in the loss, and the learning rate decides how far to step in the opposite direction.
It appears in every gradient-based method: in neural networks, logistic regression, the training of large language models, and in tree boosting under the name "eta" or "shrinkage". It is usually the first hyperparameter worth tuning.
Intuition: you are walking down a mountain in the dark. Steps that are too small — you will get there, but not before morning. Steps that are too long — you overshoot the valley and land higher up on the opposite slope, and each further jump takes you higher still.
Mechanism — why it works this way
The gradient is local information: it tells you how the loss changes in the immediate neighbourhood of the current weights, but nothing about how far away the bottom of the valley lies. By a Taylor expansion, the loss after a step is approximately L(w) − η‖∇L‖², so for a sufficiently small η the loss must decrease. The approximation breaks down when the step is too long relative to the curvature.
For a quadratic loss whose largest curvature is L (the largest eigenvalue of the Hessian), plain gradient descent converges exactly when η < 2/L, and converges fastest along the steepest direction at η ≈ 1/L. With η > 2/L every step overshoots the bottom and lands higher on the opposite slope, and the distance from the minimum grows geometrically — this is divergence. Every smooth loss looks locally like such a parabola, hence the practical rule "divide the learning rate by 2–10 when the loss jumps up".
With a constant step, stochastic gradient descent (computed on a mini-batch) never reaches the very bottom: gradient noise leaves the weights circling around the minimum within a radius proportional to η. The Robbins–Monro conditions say when a decreasing step converges: the sum of the steps must be infinite (so we can reach anywhere), and the sum of their squares finite (so the noise eventually dies down). This is the rationale for learning rate schedules: start with a larger step to descend quickly, then reduce it to settle at the bottom.
A caveat: curvature is not the same in all directions. A single learning rate is then simultaneously too large for the steep directions and too small for the flat ones — the main reason inputs are standardized and adaptive optimizers such as Adam are used.
By example
On the Diabetes dataset (442 patients, 10 standardized features) I fitted linear regression with plain gradient descent in numpy, using the loss (1/2n)‖Xw − y‖². The loss is exactly quadratic, so the threshold can be computed: the largest eigenvalue of the Hessian XᵀX/n is 4.02, so 2/L = 0.497. After 200 steps at η = 0.47 (just below the threshold) the mean squared error is 2,864, almost the least-squares optimum (2,860). At η = 0.52 (just above the threshold) it is 15,180 after 10 steps, 2.5·10⁷ after 50, and 6.6·10¹⁹ after 200. A 10% difference in learning rate separates convergence from explosion. At the other end, η = 0.01 leaves an error of 3,324 after 200 steps: the direction is right, it is just too slow.
A neural network behaves similarly. MLPClassifier (64 hidden neurons, Adam, 30 epochs, random_state=0) on Digits 8×8 scores 44.9% on the test set at lr = 0.0001 (too slow), 94.2% at 0.001, 97.1% at 0.01 and 18.2% at 1.0 (training fell apart).
In practice
- Typical starting points: SGD 0.01–0.1; Adam/AdamW 0.001 (the defaults
lr=1e-3in PyTorch andlearning_rate_init=0.001inMLPClassifier). In XGBoostetadefaults to 0.3, in LightGBMlearning_rateto 0.1. - Search the learning rate on a logarithmic scale (0.1 → 0.01 → 0.001), not a linear one.
- Schedules: step decay, cosine, cyclical; a warm-up from a small η protects against blow-ups at the start. In LLM training the standard is warm-up plus cosine, with a peak value of the order of 10⁻⁴.
- The linear scaling rule: a k times larger batch → a k times larger learning rate (works up to a certain k).
- Common mistake: the loss rises or becomes NaN — the learning rate is too large; the loss falls slowly and linearly over many epochs — it is too small. Both look like "the model isn't learning".
Frequently asked questions
- What learning rate should I start with?
- For Adam 0.001, for SGD with momentum 0.01–0.1 — then try an order of magnitude above and below. A range test (Smith 2017) is useful: increase the learning rate gradually during a short training run and pick a value just before the one at which the loss starts to rise.
- What happens when the learning rate is too large or too small?
- Too large: the step overshoots the bottom of the valley, the loss oscillates or rises, and in the extreme the weights run off to infinity (NaN). Too small: training progresses very slowly, may get stuck on a plateau and look as if it is not learning, even though the direction is right.
- Why is the learning rate reduced during training?
- At the start, a large step quickly brings the weights close to a good region. Near the bottom, mini-batch gradient noise makes a constant step circle around the minimum instead of settling into it. Reducing the step lets that noise die down — the intuition behind the Robbins–Monro conditions.
Sources
- Goodfellow, Bengio, Courville (2016). Deep Learning, ch. 8.3.1 "Stochastic gradient descent", 11.4.1 "Manual hyperparameter tuning". https://www.deeplearningbook.org/contents/optimization.html
- Zhang et al. Dive into Deep Learning, ch. 12.3 "Gradient descent", 12.11 "Learning rate scheduling". https://d2l.ai/chapter_optimization/gd.html
- Robbins, H., Monro, S. (1951). "A stochastic approximation method". Annals of Mathematical Statistics 22(3), 400–407.
- Smith, L. (2017). "Cyclical learning rates for training neural networks". WACV. arXiv:1506.01186
- Goyal et al. (2017). "Accurate, large minibatch SGD: training ImageNet in 1 hour". arXiv:1706.02677