06 · Neural nets · 4 min read · Interactive · updated
What is regularization and how do L2 and dropout reduce overfitting?
In short
Regularization stops a model from fitting noise: a penalty on large weights (L2), randomly switching off neurons (dropout) and ending training early.
What it is
Regularization is any modification of training intended to reduce error on new data at the cost of (usually) slightly higher training error. The three most common in neural networks are an L2 penalty on large weights (weight decay), dropout — randomly zeroing neurons during training — and early stopping.
In linear models, regularization means ridge (L2) and lasso (L1); in trees, limiting depth and the λ and γ penalties in XGBoost; in large language models, weight decay and dropout during fine-tuning.
Intuition: a student who is not allowed to bring long crib sheets has to remember principles rather than specific problems. Regularization limits the "space on the crib sheet": the model can fit the data, but only with a simple, smooth solution.
Mechanism — why it works this way
L2: we add λ/2 × the sum of squared weights to the loss. The gradient of this penalty is λ × weight, so each step pulls the weights towards zero in proportion to their size: weight ← weight − η(gradient + λ × weight). Large weights mean sharp reactions to small changes in the input — exactly what allows a single example to be memorized. The penalty favours solutions with small weights, i.e. smoother functions; in Bayesian terms it is a normal prior on the weights. An L1 penalty (the sum of absolute values) instead sets some weights to zero — it performs feature selection.
Dropout (Srivastava et al. 2014): at every step each hidden neuron is switched off with probability p (typically 0.5 for dense layers, 0.1–0.2 in transformers). A neuron cannot rely on the presence of any particular neighbour, because that neighbour is absent in some of the steps — it has to learn features that are useful in many combinations. Equivalently, we train exponentially many weight-sharing subnetworks at once and average them approximately at prediction time (all neurons on, activations rescaled). Dropout is turned off for evaluation.
Early stopping ends training when the validation loss stops falling. It acts as regularization because the number of steps limits how far the weights can move from their starting point; for linear models it is mathematically close to L2 (Goodfellow et al. 2016).
A caveat: regularization needs tuning. Too large a λ or p means underfitting — the model fails to fit even the true rules. The strength is chosen on validation data, never on training data, because on the training data regularization always "hurts".
By example
I split the Diabetes dataset (442 patients, 10 features, target: disease progression after one year) 75/25 (random_state=0) and expanded the features with all squares and pairwise products (PolynomialFeatures(2)), which gave 65 features for 331 training examples. I then fitted ridge regression (Ridge, i.e. an L2 penalty) with different strengths α, after standardization.
With almost no penalty (α = 10⁻⁶) the model has R² = 0.647 on training but only 0.244 on test, with a weight-vector norm of 1,139 — it has fitted the noise. At α = 100 the training R² drops to 0.603, but the test R² rises to 0.340, and the weight norm shrinks to 37. At α = 10,000 the penalty is too strong: R² is 0.113 on training and 0.080 on test (underfitting). For comparison, plain linear regression on the 10 original features scores 0.359 on test: the richer model pays off only when it is well regularized.
In practice
- scikit-learn:
LogisticRegression(C=1.0)— C is the inverse of λ, so a smaller C means stronger regularization;Ridge(alpha=...),Lasso(alpha=...);MLPClassifier(alpha=1e-4)is an L2 penalty. - PyTorch:
weight_decayin the optimizer (true decay in AdamW, an L2 penalty added to the gradient in Adam);nn.Dropout(p=0.5), and remembermodel.eval()for evaluation. - XGBoost:
reg_lambda=1(L2 on leaf values),reg_alpha=0,gamma=0,max_depth=6,min_child_weight=1. - Transformers and LLMs: dropout 0.1, weight decay 0.01–0.1; in large models trained for a single epoch dropout is often omitted.
- Common mistake: evaluating a model with dropout without switching to evaluation mode — predictions change randomly between calls.
Frequently asked questions
- How does L1 regularization differ from L2?
- L2 penalizes squared weights: it shrinks all of them proportionally, sets none to zero and gives smooth solutions. L1 penalizes absolute values: it shrinks all of them by a constant amount, so small weights land exactly at zero — the model performs feature selection. Elastic net combines both penalties.
- How does dropout work and why does it help?
- At every training step it randomly zeroes some of the neurons (usually half in dense layers). No neuron can rely on particular neighbours, so the network learns features that are robust to their absence. At prediction time all neurons are on — an approximate average over many subnetworks.
- When should I use regularization?
- When the training score is clearly better than the validation score and more data is not available. Start with early stopping and weight decay (λ of the order of 10⁻⁴–10⁻²), then add dropout. If both scores are poor, the problem is underfitting — regularization will make it worse.
Sources
- Srivastava, Hinton, Krizhevsky, Sutskever, Salakhutdinov (2014). "Dropout: a simple way to prevent neural networks from overfitting". JMLR 15, 1929–1958.
- Krogh, A., Hertz, J. (1992). "A simple weight decay can improve generalization". NeurIPS 4.
- Goodfellow, Bengio, Courville (2016). Deep Learning, ch. 7.1 "Parameter norm penalties", 7.8 "Early stopping", 7.12 "Dropout". https://www.deeplearningbook.org/contents/regularization.html
- Hastie, Tibshirani, Friedman (2009). The Elements of Statistical Learning, 2nd ed., ch. 3.4 "Shrinkage methods".
- Zhang et al. Dive into Deep Learning, ch. 3.7 "Weight decay", 5.6 "Dropout". https://d2l.ai/chapter_multilayer-perceptrons/dropout.html