06 · Neural nets · 5 min read · Interactive · updated
What is dropout in neural networks and why does it prevent overfitting?
In short
During training, dropout randomly switches off some neurons, so the network cannot rely on individual connections and generalizes better to new data.
What it is
Dropout is a regularization technique for neural networks in which, during training, every neuron in a given layer is independently switched off (its output set to zero) at every step with probability p, typically between 0.1 and 0.5. At prediction time all neurons are active. As a result the network cannot base its decisions on a few specialized neurons and has to spread its knowledge across many independent paths.
Intuition: a team in which, every day, a random half of the people are on holiday. Nobody can be the only person who knows an important process, because tomorrow they may not be there. Knowledge has to be duplicated and understandable to many. Such a team is more robust to surprises, even though on any given day it works a little more slowly.
Dropout was proposed by Hinton and colleagues in 2012, and the full description was published by Srivastava et al. in 2014. The technique was a key ingredient of the first big successes of convolutional networks in image recognition. Today transformers use it in smaller doses, and in the largest language models it is sometimes switched off entirely, because with huge amounts of data overfitting is less of a problem.
Mechanism — why it works this way
An overfit network often develops interdependencies between neurons (co-adaptation): neuron A only makes sense together with a specific error of neuron B that corrects it. Such arrangements fit the training set well but are fragile. When any neuron can vanish at any step, these arrangements stop paying off. Each neuron has to detect something useful on its own.
A second view: dropout is a cheap ensemble of models. Each dropout mask defines a different "subnetwork", and a network with n neurons has 2ⁿ possible subnetworks that share weights. Training with dropout trains a vast number of these subnetworks at once, one step for each. At prediction time the full network approximates the average of all of them. Averaging many models reduces variance, which is a classic route to better generalization.
For the averaging to work out, the scale of the activations must be the same in training and in prediction. Modern libraries use inverted dropout: during training the remaining neurons are multiplied by 1/(1 − p). The expected value of every activation is then the same as without dropout, and nothing needs to change at prediction time.
Dropout adds noise to the gradient, so training takes longer and the training loss is higher and more jagged. That is intended: regularization deliberately makes it harder to fit the training set. Too large a p leads to underfitting, and in small networks dropout can take away capacity the model needs.
Limitations: dropout works poorly with batch normalization, because it changes the activation statistics between training and prediction. In convolutional networks, zeroing individual pixels of feature maps achieves little, because neighbouring pixels carry almost the same information, so variants that zero whole channels are used instead.
By example
A layer has four activations: 0.8, 0.2, 0.5, 1.0, which sum to 2.5. With p = 0.5 the mask drawn is (1, 0, 0, 1). After applying inverted dropout the values are 1.6, 0, 0, 2.0, which sum to 3.6. A single step is therefore very noisy, but the average over all 16 possible masks is exactly 2.5: in expectation the signal does not change.
On the Digits 8×8 dataset, only 200 training images were used with a network of two hidden layers of 512 neurons each, about 300,000 parameters — clearly too many for that amount of data. After 500 Adam steps every version classifies the training set without error. On the 450 test images, the network without dropout averages 93.8% (five seeds, 93.3–94.2%), with p = 0.2 it reaches 94.1%, and with p = 0.5 94.6% (94.0–95.1%). The test loss falls from 0.37 to 0.30. The effect is real but modest: about four fewer errors out of 450. Dropout is a help, not a miracle cure for a sample that is too small.
In practice
- PyTorch:
nn.Dropout(p=0.5)after the activation in dense layers;nn.Dropout2dzeroes whole channels in convolutional networks. scikit-learn'sMLPClassifierdoes not offer dropout; it has L2 regularization (alpha) instead. - Always switch between
model.train()andmodel.eval(); dropout left on in evaluation mode is the most common bug and the reason for "random" predictions. - Typical values: 0.5 in large dense layers, 0.1–0.3 in transformers and blocks with normalization, 0.1–0.2 on the input, if at all.
- Dropout deliberately left on at prediction time (Monte Carlo dropout) gives a spread of outputs that can be treated as an approximate measure of model uncertainty.
- First check whether the model overfits at all (a gap between training and validation loss); if it does not, dropout will only slow training down.
Frequently asked questions
- What value of p should I choose?
- For dense layers the classic starting point is 0.5; for modern architectures it is more like 0.1. Treat p like any other hyperparameter and tune it on a validation set. Note: in PyTorch p is the probability of dropping a unit, while some descriptions give the probability of keeping it.
- Why is the training loss with dropout higher than the validation loss?
- Because during training the network runs handicapped, and during validation at full strength. This is normal and does not indicate a bug. To diagnose overfitting, track the validation loss over time rather than comparing it directly with the training loss.
- Does dropout replace other forms of regularization?
- No. It combines well with L2 weight regularization, early stopping and data augmentation. None of these techniques, however, is a substitute for more data.
Sources
- Srivastava N., Hinton G., Krizhevsky A., Sutskever I., Salakhutdinov R., "Dropout: A Simple Way to Prevent Neural Networks from Overfitting", Journal of Machine Learning Research 15, 2014, pp. 1929–1958.
- Hinton G. E. et al., "Improving neural networks by preventing co-adaptation of feature detectors", arXiv:1207.0580, 2012.
- Gal Y., Ghahramani Z., "Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning", ICML 2016.
- Goodfellow I., Bengio Y., Courville A., "Deep Learning", MIT Press, 2016, section 7.12 "Dropout".
- PyTorch documentation,
torch.nn.Dropout: https://pytorch.org/docs/stable/generated/torch.nn.Dropout.html