Theory becomes skill only when you can implement it. PyTorch has become the most widely used framework in deep-learning research and a major one in industry, thanks to its Pythonic, define-by-run design. This lecture is a hands-on tour. By the end you will have a reusable training template that avoids the most common mistakes.
Tensors and devices#
import torch
x = torch.tensor([[1., 2.], [3., 4.]])
y = torch.randn(2, 2)
print(x @ y, x.shape, x.dtype)
device = "cuda" if torch.cuda.is_available() else "cpu"
x = x.to(device) # move data to GPU if available
print(x.device)
a = torch.arange(6).reshape(2, 3)
print(a.float().mean(dim=1, keepdim=True)) # reductions, broadcasting โ like NumPy
print(torch.from_numpy(a.numpy())) # zero-copy bridge to NumPy (on CPU)All operands of an operation must live on the same device; mixing CPU and GPU tensors raises an error.
Autograd#
Set requires_grad=True and PyTorch records operations to compute gradients:
w = torch.tensor(2.0, requires_grad=True)
loss = (3 * w - 1) ** 2
loss.backward()
print(w.grad) # d/dw (3w-1)^2 = 6(3w-1) = 30Remember: gradients accumulate in .grad, so zero them each iteration; use torch.no_grad() during evaluation.
Defining models with nn.Module#
import torch.nn as nn
class MLP(nn.Module):
def __init__(self, d_in, d_hidden, d_out, p_drop=0.2):
super().__init__()
self.net = nn.Sequential(
nn.Linear(d_in, d_hidden), nn.ReLU(), nn.Dropout(p_drop),
nn.Linear(d_hidden, d_hidden), nn.ReLU(), nn.Dropout(p_drop),
nn.Linear(d_hidden, d_out))
def forward(self, x):
return self.net(x)Sub-modules assigned as attributes are registered automatically, so model.parameters(), model.to(device), model.train(), model.eval() and model.state_dict() all work.
Datasets and DataLoaders#
from torch.utils.data import Dataset, DataLoader
class TabularDataset(Dataset):
def __init__(self, X, y):
self.X = torch.as_tensor(X, dtype=torch.float32)
self.y = torch.as_tensor(y, dtype=torch.long)
def __len__(self):
return len(self.X)
def __getitem__(self, i):
return self.X[i], self.y[i]DataLoader handles batching, shuffling and parallel loading (num_workers), and pin_memory=True speeds up CPU-to-GPU transfer.
A complete, correct training template#
import torch, torch.nn as nn, torch.nn.functional as F
from torch.utils.data import DataLoader, random_split
from sklearn.datasets import make_classification
torch.manual_seed(0)
X, y = make_classification(n_samples=5000, n_features=20, n_informative=10, n_classes=3, random_state=0)
full = TabularDataset(X, y)
train_ds, val_ds = random_split(full, [4000, 1000], generator=torch.Generator().manual_seed(0))
train_dl = DataLoader(train_ds, batch_size=64, shuffle=True)
val_dl = DataLoader(val_ds, batch_size=256)
device = "cuda" if torch.cuda.is_available() else "cpu"
model = MLP(20, 128, 3).to(device)
opt = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=1e-2)
sched = torch.optim.lr_scheduler.CosineAnnealingLR(opt, T_max=30)
def run_epoch(loader, train):
model.train(train) # sets dropout/BN behaviour
total, correct, loss_sum = 0, 0, 0.0
with torch.set_grad_enabled(train):
for xb, yb in loader:
xb, yb = xb.to(device), yb.to(device)
logits = model(xb)
loss = F.cross_entropy(logits, yb) # logits, not softmax!
if train:
opt.zero_grad(set_to_none=True)
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step()
loss_sum += loss.item() * len(yb)
correct += (logits.argmax(1) == yb).sum().item()
total += len(yb)
return loss_sum / total, correct / total
best = 0.0
for epoch in range(30):
tr_loss, tr_acc = run_epoch(train_dl, train=True)
va_loss, va_acc = run_epoch(val_dl, train=False)
sched.step()
if va_acc > best:
best = va_acc
torch.save({"model": model.state_dict(), "epoch": epoch}, "best.pt")
if epoch % 5 == 0:
print(f"epoch {epoch:2d} train {tr_loss:.3f}/{tr_acc:.3f} val {va_loss:.3f}/{va_acc:.3f}")
print("best val acc:", round(best, 3))The five steps of every training iteration#
- Forward:
logits = model(xb) - Loss:
loss = criterion(logits, yb) - Zero gradients:
opt.zero_grad() - Backward:
loss.backward() - Update:
opt.step()
Checklist of common mistakes#
Saving, loading and deploying#
Save state_dict() (weights), not the whole model object. For deployment, export with torch.export / TorchScript or to ONNX, and use torch.compile(model) to speed up training and inference with graph compilation.
Ecosystem#
- torchvision, torchaudio, torchtext โ datasets and pretrained models.
- Hugging Face Transformers and Datasets โ pretrained NLP, vision and audio models.
- PyTorch Lightning / Accelerate โ reduce boilerplate, handle multi-GPU and mixed precision.
- TorchMetrics โ correct metrics across devices.