06 · Neural nets · 5 min read · Interactive · updated
How do SGD, momentum and Adam differ, and which optimizer should I choose?
In short
An optimizer turns a gradient into a step. SGD follows the gradient, momentum adds inertia, and Adam scales each weight's step by its typical gradient.
What it is
An optimizer is the algorithm that decides, from the gradient (and the history of gradients), how to change the weights in a single step. SGD takes a step proportional to the current gradient; momentum adds an exponentially decaying average of previous steps; Adam (Kingma and Ba 2014) combines momentum with a separate step scaling for every weight.
Optimizers are part of training every neural network; for large language models the standard is AdamW.
Intuition: SGD is a hiker who looks only at their feet at every step. Momentum is a ball rolling down the slope — it builds up speed on long descents and does not bounce off every bump. Adam is a hiker who also adapts their stride to the terrain: short steps on steep ground, long strides on the flat.
Mechanism — why it works this way
SGD: w ← w − η g, where g is the gradient from a mini-batch. The problem: the loss landscape has different curvature in different directions. In steep directions the same η is too large (zigzagging); in flat ones it is too small (crawling).
Momentum: v ← β v + g, w ← w − η v, typically with β = 0.9. The velocity v averages the gradients from roughly the last 1/(1 − β) = 10 steps. In directions where the gradient changes sign (zigzagging in a narrow valley), the components cancel out; in consistent directions they add up to as much as η g/(1 − β), i.e. a step up to 10 times longer. The inertia damps the zigzags and speeds things up in ravines.
Adam keeps two averages — m (the mean of the gradients, like momentum, β₁ = 0.9) and v (the mean of the squared gradients, β₂ = 0.999) — and takes the step w ← w − η m/(√v + ε). Dividing by the root mean square of the gradient means every weight gets a step of similar size regardless of the scale of its gradient: weights of features with large values (large gradients) get a smaller effective step, weights of features with small values a larger one. Adam therefore partly equalizes feature scales on its own, which is why its advantage over SGD grows when inputs are not standardized. It is no substitute for standardization: it equalizes the size of the step, not the shape of the valley. Bias correction adjusts m and v early in training, while the averages are still close to zero.
A caveat: on images Adam sometimes generalizes worse than well-tuned SGD with momentum, and an L2 penalty added to the gradient behaves differently in Adam than true weight decay — hence AdamW (Loshchilov and Hutter 2019), which subtracts the weight decay from the weights directly.
By example
On Digits 8×8 (1,347 training images, 450 test images, random_state=0) I trained an MLPClassifier with 64 hidden neurons for 20 epochs with a batch size of 64, changing only the optimizer. Plain SGD (learning rate 0.01) ends with a training loss of 1.41 and test accuracy of 84.4% — heading the right way, but slowly. The same SGD with momentum 0.9 gets down to a loss of 0.16 and an accuracy of 95.8%. Adam with the default lr = 0.001 gives 95.6%, and with lr = 0.01 98.2%.
Then I multiplied the pixels by 160, so that inputs ranged up to 160 instead of up to 1. SGD — with and without momentum — stopped learning: 10.0–10.4% accuracy, i.e. guessing among 10 digits. Adam with lr = 0.001 still reached 91.3%, and with lr = 0.01 94.4%: adaptive step scaling rescued training, although the result is still worse than on properly scaled data.
In practice
- PyTorch defaults:
Adam(lr=1e-3, betas=(0.9, 0.999), eps=1e-8);SGD(lr, momentum=0)— momentum must be switched on explicitly (usually 0.9). In scikit-learnMLPClassifier(solver='adam')is the default. - LLMs: AdamW with β₂ = 0.95 instead of 0.999 for stability, weight decay around 0.1, a learning rate of the order of 10⁻⁴ with warm-up and a cosine schedule.
- First choice for a new problem: Adam or AdamW with lr = 0.001. SGD with momentum and a schedule when the last hundredth of a percent on images matters.
- Adam needs two extra buffers the size of the model — for large models that is a significant memory cost.
- Common mistake: carrying the learning rate over from SGD (0.1) to Adam — for Adam that is usually two orders of magnitude too large.
Frequently asked questions
- Adam or SGD — which should I choose?
- Start with Adam (or AdamW): it is less sensitive to the learning rate and to feature scale, and works well out of the box. SGD with momentum and a good schedule can generalize better in computer vision, but needs more tuning. For LLMs the standard is AdamW.
- What does momentum do in an optimizer?
- It maintains a "velocity" — a decaying average of previous gradients — and adds it to the step. In directions where the gradient oscillates, the oscillations cancel out; in consistent directions the step grows up to tenfold (with β = 0.9). Training moves more smoothly and faster through narrow valleys.
- How is Adam different from AdamW?
- In Adam the L2 penalty is added to the gradient and then divided by √v, so weights with large gradients are regularized more weakly. AdamW subtracts the weight decay from the weights separately, outside the adaptive scaling, which gives consistent regularization and better results — which is why it is the standard.
Sources
- Kingma, D., Ba, J. (2014). "Adam: a method for stochastic optimization". ICLR 2015. arXiv:1412.6980
- Loshchilov, I., Hutter, F. (2019). "Decoupled weight decay regularization". ICLR. arXiv:1711.05101
- Sutskever, Martens, Dahl, Hinton (2013). "On the importance of initialization and momentum in deep learning". ICML.
- Goodfellow, Bengio, Courville (2016). Deep Learning, ch. 8.3 "Basic algorithms", 8.5 "Algorithms with adaptive learning rates". https://www.deeplearningbook.org/contents/optimization.html
- Zhang et al. Dive into Deep Learning, ch. 12.6 "Momentum", 12.10 "Adam". https://d2l.ai/chapter_optimization/adam.html