Once a model is trained and validated, it must be served: made available to produce predictions for applications, dashboards or users. The right serving pattern depends on how fresh predictions must be, how many are needed, and how quickly. Choosing well saves cost and complexity; choosing poorly can make a good model useless.
Serving patterns#
| Pattern | How | Latency | Example |
|---|---|---|---|
| Batch | Score many records on a schedule; store results | Minutes–hours (predictions precomputed) | Nightly risk scores for all registered households |
| Online (real-time) API | Request–response over HTTP/gRPC | Milliseconds–seconds | Classify a message when it arrives |
| Streaming | Consume events from a queue (e.g. Kafka), emit predictions | Seconds | Fraud detection on transactions |
| Edge / on-device | Model runs on phone, browser or device | Local, offline | Crop-disease detection in the field |
Prefer batch when predictions are not needed instantly: it is simpler, cheaper, easier to monitor and retry. Use online serving only when fresh, per-request predictions are required.
Building an online prediction API with FastAPI#
# serve.py — run with: uvicorn serve:app --host 0.0.0.0 --port 8000
import time, logging, joblib
import pandas as pd
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
MODEL_VERSION = "triage-classifier:7"
model = joblib.load("artifacts/model.joblib") # load ONCE at startup, not per request
app = FastAPI(title="Triage model API", version=MODEL_VERSION)
log = logging.getLogger("predictions")
class Household(BaseModel): # input validation with types and ranges
members: int = Field(ge=1, le=30)
children_under_5: int = Field(ge=0, le=15)
monthly_income: float = Field(ge=0)
has_disability: bool
district: str
class Prediction(BaseModel):
priority_score: float
priority: str
model_version: str
@app.get("/health")
def health():
return {"status": "ok", "model_version": MODEL_VERSION}
@app.post("/predict", response_model=Prediction)
def predict(h: Household):
t0 = time.perf_counter()
try:
X = pd.DataFrame([h.model_dump()])
score = float(model.predict_proba(X)[0, 1])
except Exception:
log.exception("prediction failed")
raise HTTPException(status_code=500, detail="Prediction failed")
priority = "high" if score >= 0.7 else "medium" if score >= 0.4 else "low"
log.info("latency_ms=%.1f score=%.3f version=%s", (time.perf_counter() - t0) * 1000, score, MODEL_VERSION)
return Prediction(priority_score=round(score, 4), priority=priority, model_version=MODEL_VERSION)Key practices illustrated:
- Load the model once at startup.
- Validate inputs with a schema (Pydantic) — reject malformed requests with clear errors rather than producing garbage predictions.
- Return the model version with every prediction for traceability.
- Health endpoints for load balancers and orchestrators.
- Structured logging of latency and outputs (never log sensitive personal data unnecessarily).
- Same feature code in training and serving — share the preprocessing pipeline (e.g. a scikit-learn
Pipelinesaved as one artefact) to avoid training–serving skew.
Latency and throughput#
- Measure p50, p95 and p99 latency — tail latency matters for user experience.
- Optimise the model (smaller architecture, quantisation, ONNX Runtime/TensorRT), batch requests (dynamic batching groups concurrent requests for GPU efficiency), cache frequent results, and keep feature lookups fast.
- Scale horizontally (more replicas behind a load balancer) and autoscale on load.
Dedicated serving frameworks#
For deep learning and larger deployments: TorchServe, TensorFlow Serving, NVIDIA Triton Inference Server (multi-framework, dynamic batching, GPU sharing), BentoML, Ray Serve, KServe on Kubernetes, and LLM-specific engines (vLLM, TGI). Managed cloud endpoints are an option when data policies allow.
Safe rollout strategies#
Never replace a production model all at once:
- Shadow deployment: the new model receives a copy of live traffic and its predictions are logged but not used — compare with the current model safely.
- Canary release: route a small fraction (e.g. 5%) of traffic to the new model; watch metrics; increase gradually.
- Blue–green deployment: run old and new environments side by side and switch traffic instantly, with instant rollback.
- A/B testing: randomised comparison of business/outcome metrics (see the A/B testing lecture).