Designing a CNN requires fluency with a few geometric knobs โ padding, stride, pooling and dilation. They determine the shape of every feature map, the computational cost, and the receptive field of each unit (how much of the input it can see). Shape mismatches are among the most common errors in CNN code, so let us master the arithmetic.
The output-size formula#
For input size $n$ (height or width), kernel size $k$, padding $p$, stride $s$ and dilation $d$:
With $d = 1$ this simplifies to $\left\lfloor\frac{n + 2p - k}{s}\right\rfloor + 1$.
| Input | Kernel | Padding | Stride | Output |
|---|---|---|---|---|
| 32 | 3 | 0 | 1 | 30 |
| 32 | 3 | 1 | 1 | 32 ("same") |
| 32 | 3 | 1 | 2 | 16 |
| 224 | 7 | 3 | 2 | 112 (ResNet stem) |
| 28 | 5 | 0 | 1 | 24 |
Padding#
Without padding ("valid" convolution), each layer shrinks the feature map by $k - 1$ and border pixels are used by fewer output units. Zero padding with $p = (k - 1)/2$ for odd $k$ keeps the size unchanged ("same" padding), making it easy to stack many layers. Alternatives such as reflection padding reduce border artefacts in image generation.
Stride#
A stride of $s$ moves the kernel $s$ pixels at a time, downsampling the output by roughly $s$ in each dimension. Strided convolutions are a learnable alternative to pooling and are common in modern architectures. Downsampling reduces computation in later layers and enlarges receptive fields.
Pooling#
Pooling summarises each local window with a fixed function:
- Max pooling โ keeps the strongest activation; provides a little translation invariance ("the feature is present somewhere in this window").
- Average pooling โ smooths.
- Global average pooling (GAP) โ averages each channel over the entire map, producing one number per channel. Introduced in Network-in-Network and used in GoogLeNet and ResNet, it replaces huge fully connected layers, drastically reducing parameters and allowing variable input sizes.
Pooling has no parameters. A $2 \times 2$ max pool with stride 2 halves height and width.
import torch
import torch.nn as nn
x = torch.randn(1, 3, 224, 224)
layers = [
("conv7 s2 p3", nn.Conv2d(3, 64, 7, stride=2, padding=3)),
("maxpool3 s2 p1", nn.MaxPool2d(3, stride=2, padding=1)),
("conv3 p1", nn.Conv2d(64, 64, 3, padding=1)),
("conv3 s2 p1", nn.Conv2d(64, 128, 3, stride=2, padding=1)),
("dilated conv3 d2 p2", nn.Conv2d(128, 128, 3, padding=2, dilation=2)),
("global avg pool", nn.AdaptiveAvgPool2d(1)),
]
for name, layer in layers:
x = layer(x)
print(f"{name:<20} -> {tuple(x.shape)}")Receptive fields#
The receptive field is the region of the input that can influence a unit. For a stack of layers with kernel sizes $k_l$ and strides $s_l$, it grows as
Strides early in the network multiply the growth of all later layers. Examples:
- Three stacked $3 \times 3$ convs with stride 1: $r = 1 + 2 + 2 + 2 = 7$.
- A $3 \times 3$ conv after a stride-2 layer adds $2 \times 2 = 4$ instead of 2.
The effective receptive field is smaller than the theoretical one: Luo et al. (2016) showed the influence of input pixels is roughly Gaussian-shaped, concentrated in the centre. Tasks that need global context (segmentation of large objects, scene understanding) therefore benefit from dilated convolutions, pyramid pooling or attention.
Dilation#
A dilated kernel inserts gaps between its taps. A $3 \times 3$ kernel with dilation 2 covers a $5 \times 5$ area with 9 weights. Stacking dilations 1, 2, 4, 8 grows the receptive field exponentially while keeping resolution โ the design of WaveNet for audio and DeepLab for segmentation. Watch for gridding artefacts if dilation rates share common factors.
Trade-offs in downsampling#
Downsampling saves computation and increases receptive field but discards spatial detail. Classification tolerates it well; dense prediction (segmentation, keypoints) needs detail back, which is why those architectures use skip connections from high-resolution layers (U-Net, FPN).
Standard downsampling can also break shift invariance through aliasing: a one-pixel input shift can change the output noticeably. Blurring before subsampling ("anti-aliased CNNs") improves consistency.