Activation functions are the non-linear heart of a neural network. Change them and you change how fast a network trains, whether gradients survive through many layers, and sometimes the final accuracy. The history of deep learning's progress is partly the history of better activations.
Sigmoid#
Outputs lie in $(0, 1)$ โ useful for probabilities in output layers and gates. As a hidden activation it has serious problems:
- Saturation: for large $|z|$, $\sigma'(z) \approx 0$ โ gradients vanish.
- Maximum derivative 0.25: backpropagating through many sigmoid layers multiplies many factors โค 0.25, shrinking gradients exponentially.
- Not zero-centred: outputs are always positive, making gradients of the next layer's weights all share a sign, which causes zig-zagging updates.
Tanh#
Zero-centred with outputs in $(-1, 1)$ and maximum derivative 1 โ better than sigmoid, but still saturates. It is still used inside LSTMs and GRUs.
ReLU: the revolution#
Popularised around 2010โ2012 (Nair & Hinton; Glorot et al.; AlexNet), ReLU transformed deep learning:
- No saturation for positive inputs: gradient is exactly 1, so it passes through many layers undiminished.
- Cheap: a comparison.
- Sparse activations: many units output exactly zero.
Its weakness is the dying ReLU problem: if a unit's pre-activation becomes negative for all inputs (e.g. after a large update), its gradient is zero forever and it never recovers.
ReLU variants#
| Activation | Formula | Notes |
|---|---|---|
| Leaky ReLU | $\max(\alpha z, z)$, $\alpha \approx 0.01$ | Small negative slope avoids dead units |
| PReLU | Leaky ReLU with learned $\alpha$ | Used in some vision models |
| ELU | $z$ if $z > 0$, else $\alpha(e^z - 1)$ | Smooth, negative values push mean activations towards 0 |
| SELU | Scaled ELU | Self-normalising networks under specific conditions |
Smooth modern activations: GELU and Swish/SiLU#
GELU (Gaussian Error Linear Unit, Hendrycks & Gimpel 2016) weights the input by the probability that a standard Gaussian is below it:
Swish / SiLU: $z\,\sigma(\beta z)$ (with $\beta = 1$ for SiLU).
Both are smooth, non-monotonic near zero, and behave like ReLU for large positive inputs. GELU is the default in BERT and GPT-style transformers; SiLU is common in vision models such as EfficientNet.
Gated linear units: SwiGLU#
Many recent large language models use a gated feed-forward layer:
One projection acts as a learned, input-dependent gate on another. Shazeer (2020) found GLU variants improved transformer quality at equal compute, and SwiGLU has since been adopted widely in open LLMs.
Output-layer activations#
The output activation must match the task and loss:
| Task | Output activation | Loss |
|---|---|---|
| Binary classification | Sigmoid (usually fused into the loss) | Binary cross-entropy with logits |
| Multiclass | Softmax (fused) | Cross-entropy |
| Multi-label | Independent sigmoids | Binary cross-entropy per label |
| Regression | None (identity) | MSE / MAE / Huber |
| Positive targets | Softplus or exp | Appropriate likelihood |
Visualising and checking gradients#
import torch
z = torch.linspace(-5, 5, 11, requires_grad=True)
acts = {"sigmoid": torch.sigmoid, "tanh": torch.tanh, "relu": torch.relu,
"gelu": torch.nn.functional.gelu, "silu": torch.nn.functional.silu}
for name, f in acts.items():
g, = torch.autograd.grad(f(z).sum(), z)
print(f"{name:<8} grad at z=-5..5: {[round(v, 2) for v in g.tolist()]}")
# Gradient surviving through 20 layers of each activation (product of derivatives at z = 1)
for name, f in [("sigmoid", torch.sigmoid), ("tanh", torch.tanh), ("relu", torch.relu)]:
x = torch.tensor(1.0, requires_grad=True); y = x
for _ in range(20):
y = f(y)
y.backward()
print(f"{name:<8} d(output)/d(input) through 20 layers: {x.grad.item():.2e}")The sigmoid chain's gradient collapses towards zero; ReLU's survives.