๐Ÿ”— Deep Learning ยท Lecture 28 of 38

PyTorch Fundamentals: Tensors, Autograd, Modules and the Training Loop

A practical tour of PyTorch โ€” tensors and devices, autograd, nn.Module, Dataset and DataLoader, optimisers, and a complete, correct training and evaluation loop you can reuse in every project.

Theory becomes skill only when you can implement it. PyTorch has become the most widely used framework in deep-learning research and a major one in industry, thanks to its Pythonic, define-by-run design. This lecture is a hands-on tour. By the end you will have a reusable training template that avoids the most common mistakes.

Tensors and devices#

python
import torch

x = torch.tensor([[1., 2.], [3., 4.]])
y = torch.randn(2, 2)
print(x @ y, x.shape, x.dtype)

device = "cuda" if torch.cuda.is_available() else "cpu"
x = x.to(device)                      # move data to GPU if available
print(x.device)

a = torch.arange(6).reshape(2, 3)
print(a.float().mean(dim=1, keepdim=True))   # reductions, broadcasting โ€” like NumPy
print(torch.from_numpy(a.numpy()))            # zero-copy bridge to NumPy (on CPU)

All operands of an operation must live on the same device; mixing CPU and GPU tensors raises an error.

Autograd#

Set requires_grad=True and PyTorch records operations to compute gradients:

python
w = torch.tensor(2.0, requires_grad=True)
loss = (3 * w - 1) ** 2
loss.backward()
print(w.grad)                         # d/dw (3w-1)^2 = 6(3w-1) = 30

Remember: gradients accumulate in .grad, so zero them each iteration; use torch.no_grad() during evaluation.

Defining models with nn.Module#

python
import torch.nn as nn

class MLP(nn.Module):
    def __init__(self, d_in, d_hidden, d_out, p_drop=0.2):
        super().__init__()
        self.net = nn.Sequential(
            nn.Linear(d_in, d_hidden), nn.ReLU(), nn.Dropout(p_drop),
            nn.Linear(d_hidden, d_hidden), nn.ReLU(), nn.Dropout(p_drop),
            nn.Linear(d_hidden, d_out))
    def forward(self, x):
        return self.net(x)

Sub-modules assigned as attributes are registered automatically, so model.parameters(), model.to(device), model.train(), model.eval() and model.state_dict() all work.

Datasets and DataLoaders#

python
from torch.utils.data import Dataset, DataLoader

class TabularDataset(Dataset):
    def __init__(self, X, y):
        self.X = torch.as_tensor(X, dtype=torch.float32)
        self.y = torch.as_tensor(y, dtype=torch.long)
    def __len__(self):
        return len(self.X)
    def __getitem__(self, i):
        return self.X[i], self.y[i]

DataLoader handles batching, shuffling and parallel loading (num_workers), and pin_memory=True speeds up CPU-to-GPU transfer.

A complete, correct training template#

python
import torch, torch.nn as nn, torch.nn.functional as F
from torch.utils.data import DataLoader, random_split
from sklearn.datasets import make_classification

torch.manual_seed(0)
X, y = make_classification(n_samples=5000, n_features=20, n_informative=10, n_classes=3, random_state=0)
full = TabularDataset(X, y)
train_ds, val_ds = random_split(full, [4000, 1000], generator=torch.Generator().manual_seed(0))
train_dl = DataLoader(train_ds, batch_size=64, shuffle=True)
val_dl = DataLoader(val_ds, batch_size=256)

device = "cuda" if torch.cuda.is_available() else "cpu"
model = MLP(20, 128, 3).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-2)
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=30)

def run_epoch(loader, train):
    model.train(train)                           # sets dropout/BN behaviour
    total, correct, loss_sum = 0, 0, 0.0
    with torch.set_grad_enabled(train):
        for xb, yb in loader:
            xb, yb = xb.to(device), yb.to(device)
            logits = model(xb)
            loss = F.cross_entropy(logits, yb)   # logits, not softmax!
            if train:
                opt.zero_grad(set_to_none=True)
                loss.backward()
                torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
                opt.step()
            loss_sum += loss.item() * len(yb)
            correct += (logits.argmax(1) == yb).sum().item()
            total += len(yb)
    return loss_sum / total, correct / total

best = 0.0
for epoch in range(30):
    tr_loss, tr_acc = run_epoch(train_dl, train=True)
    va_loss, va_acc = run_epoch(val_dl, train=False)
    sched.step()
    if va_acc > best:
        best = va_acc
        torch.save({"model": model.state_dict(), "epoch": epoch}, "best.pt")
    if epoch % 5 == 0:
        print(f"epoch {epoch:2d}  train {tr_loss:.3f}/{tr_acc:.3f}  val {va_loss:.3f}/{va_acc:.3f}")
print("best val acc:", round(best, 3))

The five steps of every training iteration#

  1. Forward: logits = model(xb)
  2. Loss: loss = criterion(logits, yb)
  3. Zero gradients: opt.zero_grad()
  4. Backward: loss.backward()
  5. Update: opt.step()

Checklist of common mistakes#

Saving, loading and deploying#

Save state_dict() (weights), not the whole model object. For deployment, export with torch.export / TorchScript or to ONNX, and use torch.compile(model) to speed up training and inference with graph compilation.

Ecosystem#

  • torchvision, torchaudio, torchtext โ€” datasets and pretrained models.
  • Hugging Face Transformers and Datasets โ€” pretrained NLP, vision and audio models.
  • PyTorch Lightning / Accelerate โ€” reduce boilerplate, handle multi-GPU and mixed precision.
  • TorchMetrics โ€” correct metrics across devices.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ”— Deep Learning

Mixed-Precision Training and GPU Efficiency

Training in 16-bit arithmetic roughly halves memory and can multiply throughput. We explain float16 and bfloat16, loss scaling, automatic mixed precision, and other practical techniques to make GPUs work harder.

Advancedโฑ 5 min#127
๐Ÿ”— Deep Learning

Computational Graphs and Automatic Differentiation

Frameworks compute gradients of arbitrary programs automatically. We compare symbolic, numerical and automatic differentiation, contrast forward and reverse mode, and build a tiny reverse-mode autodiff engine.

Intermediateโฑ 6 min#103
๐Ÿ”— Deep Learning

Transfer Learning and Fine-Tuning

Pretrained models let you achieve strong results with small datasets. We compare feature extraction and fine-tuning, explain discriminative learning rates and layer freezing, and discuss when transfer helps or hurts.

Intermediateโฑ 5 min#123