ML Atlas

06 · Neural nets · 4 min read · Interactive · updated

What is an activation function and why did ReLU replace the sigmoid?

In short

An activation function is the nonlinearity applied to a neuron's weighted sum. Without it a deep network is linear, and its shape decides how gradients flow.

What it is

An activation function is a nonlinear transformation applied to a neuron's weighted sum before the result is passed to the next layer. The most common ones are the sigmoid (squashes values into the range 0–1), tanh (from −1 to 1), ReLU (zero for negative inputs, identity for positive ones) and its variants: Leaky ReLU, ELU, GELU, SiLU.

An activation function sits in every hidden layer of a neural network; in transformers and large language models the standard choices are GELU or SwiGLU.

Intuition: purely linear layers are like folding a sheet of paper along straight lines that immediately unfold again — after any number of folds the sheet is flat once more. An activation is a crease that stays. Only from many such creases can you shape an arbitrary form.

Mechanism — why it works this way

The first role is nonlinearity. A composition of linear layers is still linear, so without activations a network of any depth would draw a single hyperplane. The activation "bends" the response, which lets neurons be combined into curves and regions. A ReLU network is a piecewise linear function: each neuron adds one kink.

The second role, less obvious, is letting the gradient through. Backpropagation multiplies the gradient by the derivative of the activation at every layer. The derivative of the sigmoid is σ(z)(1 − σ(z)), at most 0.25 (at zero) and practically zero for |z| > 5. With L sigmoid layers this factor alone scales the gradient by at most 0.25ᴸ — with six layers that is at least 4,000 times smaller. This is the vanishing gradient: early layers receive almost no signal and stop learning. Tanh has a derivative of up to 1 at zero, so it does better, but it too saturates at the extremes.

ReLU has a derivative of exactly 1 for positive inputs, so it does not dampen the gradient in active neurons; it is also cheap to compute and produces sparse activations. The price: for negative inputs the derivative is 0, so a neuron whose sum is negative for every example stops learning for good (dying ReLU). Leaky ReLU adds a small slope (typically 0.01) on the left to soften this; GELU and SiLU are smooth versions.

A caveat: the gradient is also multiplied by the weights, so ReLU does not guarantee stability — you need a matching initialization (He) and often normalization. At the output, the activation depends on the task: sigmoid or softmax for probabilities, none at all for regression.

By example

On Digits 8×8 (1,347 training images, 450 test images, random_state=0, pixels divided by 16) I compared MLPClassifier with different activations. With one hidden layer of 100 neurons the differences are small: no activation (identity, i.e. a linear model) 96.0% on the test set, sigmoid 97.6%, tanh 96.9%, ReLU 98.0%. The sigmoid, however, needed 596 epochs to converge, ReLU 295.

Depth changes the picture. A network with eight hidden layers of 32 neurons each (Adam, lr = 0.001) using the sigmoid did not learn at all: 10.2% accuracy and a loss of 2.303 = ln 10, i.e. uniform guessing among 10 digits — the gradient never reached the early layers. The same network with tanh reached 96.9%, and with ReLU 93.1%.

In practice

  • Default for hidden layers: ReLU (MLPClassifier(activation='relu'), nn.ReLU). Transformers: GELU (nn.GELU), SwiGLU in newer LLMs.
  • Sigmoid and tanh survive at the output, in LSTM and GRU gates, and wherever a value must stay within a range.
  • Pair ReLU with He initialization (kaiming_normal_); sigmoid and tanh with Xavier.
  • Monitor the share of neurons that always output zero (dead ReLUs); a few percent is normal, 30–50% is a problem — lower the learning rate or switch to Leaky ReLU.
  • Common mistake: a sigmoid in a deep network without normalization — training "stalls" even though the code is correct.

Frequently asked questions

Why is ReLU better than the sigmoid?
The sigmoid's derivative is at most 0.25, so the gradient shrinks at least fourfold at every layer and vanishes in a deep network. ReLU has a derivative of 1 for positive inputs, is cheap and gives sparse activations. The sigmoid has stayed at the output, where a probability is needed.
What is the vanishing gradient?
A situation in which the gradient for the early layers is close to zero because it has been multiplied by many small activation derivatives (and small weights). The early layers stop learning. Remedies: ReLU, He or Xavier initialization, batch norm or layer norm, residual connections.
Can a network without activation functions learn anything?
Only a linear function. A product of matrices is a single matrix, so a multilayer network without activations is linear or logistic regression with oddly parameterized weights. It will not learn XOR or any boundary with a kink.

Sources

  • Goodfellow, Bengio, Courville (2016). Deep Learning, ch. 6.3 "Hidden units". https://www.deeplearningbook.org/contents/mlp.html
  • Nair, V., Hinton, G. (2010). "Rectified linear units improve restricted Boltzmann machines". ICML.
  • Glorot, X., Bordes, A., Bengio, Y. (2011). "Deep sparse rectifier neural networks". AISTATS, PMLR 15.
  • Hendrycks, D., Gimpel, K. (2016). "Gaussian error linear units (GELUs)". arXiv:1606.08415
  • Zhang et al. Dive into Deep Learning, ch. 5.1.2 "Activation functions". https://d2l.ai/chapter_multilayer-perceptrons/mlp.html

See also