๐Ÿ”— Deep Learning ยท Lecture 4 of 38

Activation Functions: Sigmoid, Tanh, ReLU, GELU, SwiGLU and Softmax

The choice of non-linearity shapes how gradients flow and how networks learn. We compare the classical and modern activations, their derivatives and failure modes, and which to use where.

Activation functions are the non-linear heart of a neural network. Change them and you change how fast a network trains, whether gradients survive through many layers, and sometimes the final accuracy. The history of deep learning's progress is partly the history of better activations.

Sigmoid#

$$ \sigma(z) = \frac{1}{1 + e^{-z}}, \qquad \sigma'(z) = \sigma(z)(1 - \sigma(z)) $$

Outputs lie in $(0, 1)$ โ€” useful for probabilities in output layers and gates. As a hidden activation it has serious problems:

  • Saturation: for large $|z|$, $\sigma'(z) \approx 0$ โ€” gradients vanish.
  • Maximum derivative 0.25: backpropagating through many sigmoid layers multiplies many factors โ‰ค 0.25, shrinking gradients exponentially.
  • Not zero-centred: outputs are always positive, making gradients of the next layer's weights all share a sign, which causes zig-zagging updates.

Tanh#

$$ \tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}} = 2\sigma(2z) - 1, \qquad \tanh'(z) = 1 - \tanh^2(z) $$

Zero-centred with outputs in $(-1, 1)$ and maximum derivative 1 โ€” better than sigmoid, but still saturates. It is still used inside LSTMs and GRUs.

ReLU: the revolution#

$$ \text{ReLU}(z) = \max(0, z), \qquad \text{ReLU}'(z) = \begin{cases} 1 & z > 0 \\ 0 & z < 0 \end{cases} $$

Popularised around 2010โ€“2012 (Nair & Hinton; Glorot et al.; AlexNet), ReLU transformed deep learning:

  • No saturation for positive inputs: gradient is exactly 1, so it passes through many layers undiminished.
  • Cheap: a comparison.
  • Sparse activations: many units output exactly zero.

Its weakness is the dying ReLU problem: if a unit's pre-activation becomes negative for all inputs (e.g. after a large update), its gradient is zero forever and it never recovers.

ReLU variants#

ActivationFormulaNotes
Leaky ReLU$\max(\alpha z, z)$, $\alpha \approx 0.01$Small negative slope avoids dead units
PReLULeaky ReLU with learned $\alpha$Used in some vision models
ELU$z$ if $z > 0$, else $\alpha(e^z - 1)$Smooth, negative values push mean activations towards 0
SELUScaled ELUSelf-normalising networks under specific conditions

Smooth modern activations: GELU and Swish/SiLU#

GELU (Gaussian Error Linear Unit, Hendrycks & Gimpel 2016) weights the input by the probability that a standard Gaussian is below it:

$$ \text{GELU}(z) = z\,\Phi(z) \approx 0.5z\left(1 + \tanh\left[\sqrt{2/\pi}\,(z + 0.044715z^3)\right]\right) $$

Swish / SiLU: $z\,\sigma(\beta z)$ (with $\beta = 1$ for SiLU).

Both are smooth, non-monotonic near zero, and behave like ReLU for large positive inputs. GELU is the default in BERT and GPT-style transformers; SiLU is common in vision models such as EfficientNet.

Gated linear units: SwiGLU#

Many recent large language models use a gated feed-forward layer:

$$ \text{SwiGLU}(\mathbf{x}) = \big(\text{Swish}(\mathbf{x}\mathbf{W}_1) \odot \mathbf{x}\mathbf{W}_2\big)\mathbf{W}_3 $$

One projection acts as a learned, input-dependent gate on another. Shazeer (2020) found GLU variants improved transformer quality at equal compute, and SwiGLU has since been adopted widely in open LLMs.

Output-layer activations#

The output activation must match the task and loss:

TaskOutput activationLoss
Binary classificationSigmoid (usually fused into the loss)Binary cross-entropy with logits
MulticlassSoftmax (fused)Cross-entropy
Multi-labelIndependent sigmoidsBinary cross-entropy per label
RegressionNone (identity)MSE / MAE / Huber
Positive targetsSoftplus or expAppropriate likelihood

Visualising and checking gradients#

python
import torch

z = torch.linspace(-5, 5, 11, requires_grad=True)
acts = {"sigmoid": torch.sigmoid, "tanh": torch.tanh, "relu": torch.relu,
        "gelu": torch.nn.functional.gelu, "silu": torch.nn.functional.silu}
for name, f in acts.items():
    g, = torch.autograd.grad(f(z).sum(), z)
    print(f"{name:<8} grad at z=-5..5: {[round(v, 2) for v in g.tolist()]}")

# Gradient surviving through 20 layers of each activation (product of derivatives at z = 1)
for name, f in [("sigmoid", torch.sigmoid), ("tanh", torch.tanh), ("relu", torch.relu)]:
    x = torch.tensor(1.0, requires_grad=True); y = x
    for _ in range(20):
        y = f(y)
    y.backward()
    print(f"{name:<8} d(output)/d(input) through 20 layers: {x.grad.item():.2e}")

The sigmoid chain's gradient collapses towards zero; ReLU's survives.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Multilayer Perceptrons and the Universal Approximation Theorem

With one hidden layer, a network can approximate any continuous function โ€” so why go deep? We state the universal approximation theorem, build intuition with bumps, and explain the efficiency advantages of depth.

Intermediateโฑ 5 min#099
๐Ÿ”— Deep Learning

Loss Functions in Deep Learning: What Are We Really Optimising?

The loss function defines what "good" means to a network. We survey regression, classification, ranking and representation-learning losses, their probabilistic meaning, and common implementation mistakes.

Beginnerโฑ 5 min#101
๐Ÿ”— Deep Learning

The Perceptron: The First Learning Machine

Rosenblatt's perceptron learned to classify by correcting its mistakes. We derive its learning rule, prove the convergence theorem, reveal its XOR limitation, and see how it foreshadowed modern networks.

Beginnerโฑ 5 min#098