⚙️ MLOps & Engineering · Lecture 6 of 15

Docker and Containers for Machine Learning

"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.

A model trained with one version of PyTorch, CUDA and NumPy may fail — or silently behave differently — on a machine with other versions. Containers package code together with its entire runtime environment — operating system libraries, Python, dependencies, configuration — so it runs identically on a laptop, a server, a cloud cluster or a partner's infrastructure. Docker is the most widely used container tool, and container skills are essential for any ML engineer.

Containers vs virtual machines#

  • A virtual machine emulates a full computer, with its own operating-system kernel — heavy (gigabytes), slow to start.
  • A container shares the host kernel and isolates processes, filesystems and networking — lightweight (megabytes to a few gigabytes), starts in seconds.

Key concepts:

  • Image: an immutable template built from a Dockerfile, composed of cached layers.
  • Container: a running instance of an image.
  • Registry: where images are stored and shared (Docker Hub, GitHub Container Registry, cloud registries).
  • Volume: persistent storage mounted into a container (for data and model files).

A Dockerfile for a prediction service#

dockerfile
# Dockerfile
FROM python:3.11-slim AS base

# System settings: no .pyc files, unbuffered logs
ENV PYTHONDONTWRITEBYTECODE=1 PYTHONUNBUFFERED=1

WORKDIR /app

# 1) Install dependencies first — this layer is cached until requirements change
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# 2) Copy application code and model artefact
COPY src/ ./src/
COPY artifacts/model.joblib ./artifacts/model.joblib

# 3) Run as a non-root user for security
RUN useradd --create-home appuser
USER appuser

EXPOSE 8000
HEALTHCHECK CMD python -c "import urllib.request; urllib.request.urlopen('http://localhost:8000/health')"
CMD ["uvicorn", "src.serve:app", "--host", "0.0.0.0", "--port", "8000"]
bash
docker build -t triage-api:1.0.0 .
docker run --rm -p 8000:8000 triage-api:1.0.0
curl -X POST localhost:8000/predict -H "Content-Type: application/json" \
     -d '{"members": 5, "children_under_5": 2, "monthly_income": 8000, "has_disability": false, "district": "Dhaka"}'

Layer caching and image size#

Order instructions from least to most frequently changing: base image → system packages → Python dependencies → application code. Then editing code does not reinstall dependencies.

Keep images small and secure:

  • Use slim base images; avoid installing build tools in the final image.
  • Use multi-stage builds: compile or install in a "builder" stage and copy only what is needed into the final stage.
  • Add a .dockerignore (exclude .git, datasets, notebooks, virtual environments, secrets).
  • Clean package caches (--no-cache-dir).

GPUs in containers#

Deep-learning images need CUDA libraries matching the framework version. Use official base images (e.g. PyTorch or NVIDIA CUDA images) and run with GPU access via the NVIDIA Container Toolkit:

bash
docker run --rm --gpus all pytorch/pytorch:latest python -c "import torch; print(torch.cuda.is_available())"

The host needs a compatible NVIDIA driver; the container brings the CUDA runtime.

Training in containers#

Containers make training jobs reproducible and portable to clusters:

bash
docker run --rm --gpus all \
  -v "$PWD/data:/app/data:ro" \          # mount data read-only
  -v "$PWD/artifacts:/app/artifacts" \   # write model outputs to the host
  -e MLFLOW_TRACKING_URI=http://mlflow:5000 \
  trainer:2.3.0 python src/train.py --config configs/train.yaml

Do not bake large datasets into images; mount them or read them from storage.

Docker Compose for local stacks#

Run multi-service setups (API + database + MLflow + monitoring) with one file:

yaml
# compose.yaml
services:
  api:
    build: .
    ports: ["8000:8000"]
    depends_on: [mlflow]
  mlflow:
    image: ghcr.io/mlflow/mlflow:latest
    command: mlflow server --host 0.0.0.0 --port 5000
    ports: ["5000:5000"]

Security essentials#

Orchestration#

In production, containers are managed by orchestrators — most commonly Kubernetes — which handle scheduling, scaling, restarts, rolling updates, service discovery and resource limits (CPU, memory, GPUs). ML platforms (Kubeflow, KServe, cloud ML services) build on these foundations. For smaller deployments, a single server with Docker Compose, or a managed container service, may be entirely sufficient.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Intermediate⏱ 4 min#245
⚙️ MLOps & Engineering

GPUs and Hardware for Machine Learning

Understanding hardware helps you train faster and cheaper. We explain why GPUs suit deep learning, the roles of memory capacity and bandwidth, precision and tensor cores, estimating requirements, and choosing between local, cloud and free resources.

Intermediate⏱ 5 min#252
⚙️ MLOps & Engineering

Model Serving: Batch, Real-Time APIs and Streaming

A model creates value only when its predictions reach people and systems. We compare batch, online and streaming serving, build a FastAPI prediction service with validation, and cover latency, scaling and safe rollout.

Intermediate⏱ 5 min#246