⚙️ MLOps & Engineering · Lecture 5 of 15

Model Serving: Batch, Real-Time APIs and Streaming

A model creates value only when its predictions reach people and systems. We compare batch, online and streaming serving, build a FastAPI prediction service with validation, and cover latency, scaling and safe rollout.

Once a model is trained and validated, it must be served: made available to produce predictions for applications, dashboards or users. The right serving pattern depends on how fresh predictions must be, how many are needed, and how quickly. Choosing well saves cost and complexity; choosing poorly can make a good model useless.

Serving patterns#

PatternHowLatencyExample
BatchScore many records on a schedule; store resultsMinutes–hours (predictions precomputed)Nightly risk scores for all registered households
Online (real-time) APIRequest–response over HTTP/gRPCMilliseconds–secondsClassify a message when it arrives
StreamingConsume events from a queue (e.g. Kafka), emit predictionsSecondsFraud detection on transactions
Edge / on-deviceModel runs on phone, browser or deviceLocal, offlineCrop-disease detection in the field

Prefer batch when predictions are not needed instantly: it is simpler, cheaper, easier to monitor and retry. Use online serving only when fresh, per-request predictions are required.

Building an online prediction API with FastAPI#

python
# serve.py — run with:  uvicorn serve:app --host 0.0.0.0 --port 8000
import time, logging, joblib
import pandas as pd
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field

MODEL_VERSION = "triage-classifier:7"
model = joblib.load("artifacts/model.joblib")          # load ONCE at startup, not per request
app = FastAPI(title="Triage model API", version=MODEL_VERSION)
log = logging.getLogger("predictions")

class Household(BaseModel):                             # input validation with types and ranges
    members: int = Field(ge=1, le=30)
    children_under_5: int = Field(ge=0, le=15)
    monthly_income: float = Field(ge=0)
    has_disability: bool
    district: str

class Prediction(BaseModel):
    priority_score: float
    priority: str
    model_version: str

@app.get("/health")
def health():
    return {"status": "ok", "model_version": MODEL_VERSION}

@app.post("/predict", response_model=Prediction)
def predict(h: Household):
    t0 = time.perf_counter()
    try:
        X = pd.DataFrame([h.model_dump()])
        score = float(model.predict_proba(X)[0, 1])
    except Exception:
        log.exception("prediction failed")
        raise HTTPException(status_code=500, detail="Prediction failed")
    priority = "high" if score >= 0.7 else "medium" if score >= 0.4 else "low"
    log.info("latency_ms=%.1f score=%.3f version=%s", (time.perf_counter() - t0) * 1000, score, MODEL_VERSION)
    return Prediction(priority_score=round(score, 4), priority=priority, model_version=MODEL_VERSION)

Key practices illustrated:

  • Load the model once at startup.
  • Validate inputs with a schema (Pydantic) — reject malformed requests with clear errors rather than producing garbage predictions.
  • Return the model version with every prediction for traceability.
  • Health endpoints for load balancers and orchestrators.
  • Structured logging of latency and outputs (never log sensitive personal data unnecessarily).
  • Same feature code in training and serving — share the preprocessing pipeline (e.g. a scikit-learn Pipeline saved as one artefact) to avoid training–serving skew.

Latency and throughput#

  • Measure p50, p95 and p99 latency — tail latency matters for user experience.
  • Optimise the model (smaller architecture, quantisation, ONNX Runtime/TensorRT), batch requests (dynamic batching groups concurrent requests for GPU efficiency), cache frequent results, and keep feature lookups fast.
  • Scale horizontally (more replicas behind a load balancer) and autoscale on load.

Dedicated serving frameworks#

For deep learning and larger deployments: TorchServe, TensorFlow Serving, NVIDIA Triton Inference Server (multi-framework, dynamic batching, GPU sharing), BentoML, Ray Serve, KServe on Kubernetes, and LLM-specific engines (vLLM, TGI). Managed cloud endpoints are an option when data policies allow.

Safe rollout strategies#

Never replace a production model all at once:

  • Shadow deployment: the new model receives a copy of live traffic and its predictions are logged but not used — compare with the current model safely.
  • Canary release: route a small fraction (e.g. 5%) of traffic to the new model; watch metrics; increase gradually.
  • Blue–green deployment: run old and new environments side by side and switch traffic instantly, with instant rollback.
  • A/B testing: randomised comparison of business/outcome metrics (see the A/B testing lecture).
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

Reproducibility in Machine Learning

Can someone else — or you, in six months — get the same result? We examine sources of non-reproducibility, from seeds and GPUs to data and environments, and practical steps for reproducible research and production.

Intermediate⏱ 4 min#245
⚙️ MLOps & Engineering

Docker and Containers for Machine Learning

"It works on my machine" is not a deployment strategy. We explain containers, write efficient Dockerfiles for ML training and serving, handle GPUs, and cover image size, security and orchestration basics.

Beginner⏱ 5 min#247
⚙️ MLOps & Engineering

Experiment Tracking: Never Lose a Result Again

ML development involves hundreds of runs with different data, code and hyperparameters. We cover what to track, how to use MLflow for runs, metrics and artefacts, the model registry, and good experiment hygiene.

Beginner⏱ 4 min#244