Beginners think machine learning is about choosing algorithms. Practitioners know that algorithm choice is a small part of the job. The rest is framing the problem, understanding the data, building evaluation you can trust, and iterating. Today we walk through the complete workflow, the same one I recommend for your final-year projects.
Step 1: Frame the problem#
- What decision will the model support? A model is useful only if it changes a decision.
- What is the target? Define it precisely. "Customer churn" โ within 30 days? 90 days?
- What is the success metric, both technical (F1, RMSE) and operational (time saved, errors avoided)?
- What is the baseline? How is the task done today, and how well?
- What are the constraints? Latency, interpretability, privacy, fairness, cost.
Step 2: Collect and understand the data#
Ask where the data comes from, how it was labelled, what time period it covers, and who is missing from it. Then perform Exploratory Data Analysis (EDA): distributions, missing values, outliers, correlations, class balance, duplicates, and leakage โ features that would not be available at prediction time.
Step 3: Split the data correctly#
Split before any fitting (including scaling or feature selection) to avoid leakage:
- Training set โ fit parameters.
- Validation set โ choose hyperparameters and models.
- Test set โ final, one-time estimate of performance.
For time-dependent data, split by time (train on the past, test on the future). For data with groups (multiple records per patient), split by group so the same patient never appears in both train and test.
Step 4: Build a simple baseline#
Start with the simplest reasonable model โ predict the majority class, predict the mean, or a logistic regression. A baseline tells you whether the problem is learnable and gives a reference point. Surprisingly often, a well-tuned simple model is hard to beat on tabular data.
Step 5: Iterate#
Improve in order of expected payoff: fix data quality issues โ engineer features โ try stronger model families โ tune hyperparameters โ ensemble. Change one thing at a time and track every experiment.
Step 6: Evaluate thoroughly#
Beyond one headline metric:
- Confusion matrix and per-class metrics.
- Performance on important subgroups (region, gender, language) โ fairness checks.
- Error analysis: read 50โ100 misclassified examples and categorise the causes.
- Calibration of predicted probabilities.
- Robustness to realistic shifts.
Step 7: Deploy and monitor#
Package the model and its preprocessing together, serve predictions, and monitor input distributions, prediction distributions and โ when labels arrive โ real-world performance. Plan for retraining. (The MLOps track covers this in depth.)
A complete example#
import numpy as np
import pandas as pd
from sklearn.datasets import fetch_openml
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.dummy import DummyClassifier
from sklearn.metrics import classification_report
# 1-2. Data: the Titanic survival dataset
df = fetch_openml("titanic", version=1, as_frame=True).frame
X = df[["pclass", "sex", "age", "sibsp", "parch", "fare", "embarked"]]
y = (df["survived"].astype(int))
# 3. Split first
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)
num = ["age", "sibsp", "parch", "fare"]
cat = ["pclass", "sex", "embarked"]
prep = ColumnTransformer([
("num", Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler())]), num),
("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
("onehot", OneHotEncoder(handle_unknown="ignore"))]), cat),
])
# 4-5. Baseline and candidates, compared by cross-validation on the training set only
candidates = {
"majority baseline": DummyClassifier(strategy="most_frequent"),
"logistic regression": LogisticRegression(max_iter=1000),
"gradient boosting": HistGradientBoostingClassifier(random_state=0),
}
for name, model in candidates.items():
pipe = Pipeline([("prep", prep), ("model", model)])
scores = cross_val_score(pipe, X_tr, y_tr, cv=5, scoring="f1")
print(f"{name:<20} F1 = {scores.mean():.3f} ยฑ {scores.std():.3f}")
# 6. Final evaluation of the chosen model on the untouched test set
best = Pipeline([("prep", prep), ("model", HistGradientBoostingClassifier(random_state=0))]).fit(X_tr, y_tr)
print(classification_report(y_te, best.predict(X_te)))Notice three professional habits: preprocessing lives inside the pipeline (so it is fitted only on training folds), we compare against a baseline, and the test set is used once.
Common workflow failures#
| Failure | Symptom | Prevention |
|---|---|---|
| Data leakage | Suspiciously high validation scores | Split first; pipelines; check feature timing |
| Wrong metric | Great metric, useless model | Frame costs of errors up front |
| Test-set reuse | Reported results do not hold in production | Validation set for decisions; test once |
| Ignoring subgroups | Harm to under-represented users | Sliced evaluation |
| No monitoring | Silent degradation | Monitor drift and outcomes |