๐Ÿ”— Deep Learning ยท Lecture 18 of 38

Padding, Stride, Pooling and Receptive Fields

The geometry of convolutional layers determines output sizes, computational cost and what each unit can see. We derive the output-size formula, compare pooling types, and compute receptive fields.

Designing a CNN requires fluency with a few geometric knobs โ€” padding, stride, pooling and dilation. They determine the shape of every feature map, the computational cost, and the receptive field of each unit (how much of the input it can see). Shape mismatches are among the most common errors in CNN code, so let us master the arithmetic.

The output-size formula#

For input size $n$ (height or width), kernel size $k$, padding $p$, stride $s$ and dilation $d$:

$$ n_{\text{out}} = \left\lfloor\frac{n + 2p - d(k - 1) - 1}{s}\right\rfloor + 1 $$

With $d = 1$ this simplifies to $\left\lfloor\frac{n + 2p - k}{s}\right\rfloor + 1$.

InputKernelPaddingStrideOutput
3230130
3231132 ("same")
3231216
224732112 (ResNet stem)
2850124

Padding#

Without padding ("valid" convolution), each layer shrinks the feature map by $k - 1$ and border pixels are used by fewer output units. Zero padding with $p = (k - 1)/2$ for odd $k$ keeps the size unchanged ("same" padding), making it easy to stack many layers. Alternatives such as reflection padding reduce border artefacts in image generation.

Stride#

A stride of $s$ moves the kernel $s$ pixels at a time, downsampling the output by roughly $s$ in each dimension. Strided convolutions are a learnable alternative to pooling and are common in modern architectures. Downsampling reduces computation in later layers and enlarges receptive fields.

Pooling#

Pooling summarises each local window with a fixed function:

  • Max pooling โ€” keeps the strongest activation; provides a little translation invariance ("the feature is present somewhere in this window").
  • Average pooling โ€” smooths.
  • Global average pooling (GAP) โ€” averages each channel over the entire map, producing one number per channel. Introduced in Network-in-Network and used in GoogLeNet and ResNet, it replaces huge fully connected layers, drastically reducing parameters and allowing variable input sizes.

Pooling has no parameters. A $2 \times 2$ max pool with stride 2 halves height and width.

python
import torch
import torch.nn as nn

x = torch.randn(1, 3, 224, 224)
layers = [
    ("conv7 s2 p3", nn.Conv2d(3, 64, 7, stride=2, padding=3)),
    ("maxpool3 s2 p1", nn.MaxPool2d(3, stride=2, padding=1)),
    ("conv3 p1", nn.Conv2d(64, 64, 3, padding=1)),
    ("conv3 s2 p1", nn.Conv2d(64, 128, 3, stride=2, padding=1)),
    ("dilated conv3 d2 p2", nn.Conv2d(128, 128, 3, padding=2, dilation=2)),
    ("global avg pool", nn.AdaptiveAvgPool2d(1)),
]
for name, layer in layers:
    x = layer(x)
    print(f"{name:<20} -> {tuple(x.shape)}")

Receptive fields#

The receptive field is the region of the input that can influence a unit. For a stack of layers with kernel sizes $k_l$ and strides $s_l$, it grows as

$$ r_l = r_{l-1} + (k_l - 1)\prod_{i=1}^{l-1}s_i, \qquad r_0 = 1 $$

Strides early in the network multiply the growth of all later layers. Examples:

  • Three stacked $3 \times 3$ convs with stride 1: $r = 1 + 2 + 2 + 2 = 7$.
  • A $3 \times 3$ conv after a stride-2 layer adds $2 \times 2 = 4$ instead of 2.

The effective receptive field is smaller than the theoretical one: Luo et al. (2016) showed the influence of input pixels is roughly Gaussian-shaped, concentrated in the centre. Tasks that need global context (segmentation of large objects, scene understanding) therefore benefit from dilated convolutions, pyramid pooling or attention.

Dilation#

A dilated kernel inserts gaps between its taps. A $3 \times 3$ kernel with dilation 2 covers a $5 \times 5$ area with 9 weights. Stacking dilations 1, 2, 4, 8 grows the receptive field exponentially while keeping resolution โ€” the design of WaveNet for audio and DeepLab for segmentation. Watch for gridding artefacts if dilation rates share common factors.

Trade-offs in downsampling#

Downsampling saves computation and increases receptive field but discards spatial detail. Classification tolerates it well; dense prediction (segmentation, keypoints) needs detail back, which is why those architectures use skip connections from high-resolution layers (U-Net, FPN).

Standard downsampling can also break shift invariance through aliasing: a one-pixel input shift can change the output noticeably. Blurring before subsampling ("anti-aliased CNNs") improves consistency.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Convolutional Neural Networks: The Core Ideas

Convolutions exploit the structure of images through local connectivity, weight sharing and translation equivariance. We define the convolution operation, count parameters, and build a CNN that learns hierarchical features.

Beginnerโฑ 5 min#113
๐Ÿ”— Deep Learning

Batch Normalisation: Faster, More Stable Training

BatchNorm normalises each feature using mini-batch statistics, then rescales it with learned parameters. We derive the forward pass, explain training-versus-inference behaviour, debate why it works, and list its pitfalls.

Intermediateโฑ 5 min#109
๐Ÿ”— Deep Learning

Recurrent Neural Networks: Modelling Sequences

Sequences need memory. RNNs carry a hidden state through time with shared weights. We define the vanilla RNN, unroll it, derive backpropagation through time, and see why long dependencies are hard.

Intermediateโฑ 5 min#115