๐Ÿ“ˆ Machine Learning ยท Lecture 3 of 47

The Machine Learning Workflow: From Problem to Deployed Model

Successful ML projects follow a disciplined process โ€” problem framing, data collection, exploration, baselines, iteration, evaluation and deployment. We walk through it with a complete scikit-learn example.

Beginners think machine learning is about choosing algorithms. Practitioners know that algorithm choice is a small part of the job. The rest is framing the problem, understanding the data, building evaluation you can trust, and iterating. Today we walk through the complete workflow, the same one I recommend for your final-year projects.

Step 1: Frame the problem#

  • What decision will the model support? A model is useful only if it changes a decision.
  • What is the target? Define it precisely. "Customer churn" โ€” within 30 days? 90 days?
  • What is the success metric, both technical (F1, RMSE) and operational (time saved, errors avoided)?
  • What is the baseline? How is the task done today, and how well?
  • What are the constraints? Latency, interpretability, privacy, fairness, cost.

Step 2: Collect and understand the data#

Ask where the data comes from, how it was labelled, what time period it covers, and who is missing from it. Then perform Exploratory Data Analysis (EDA): distributions, missing values, outliers, correlations, class balance, duplicates, and leakage โ€” features that would not be available at prediction time.

Step 3: Split the data correctly#

Split before any fitting (including scaling or feature selection) to avoid leakage:

  • Training set โ€” fit parameters.
  • Validation set โ€” choose hyperparameters and models.
  • Test set โ€” final, one-time estimate of performance.

For time-dependent data, split by time (train on the past, test on the future). For data with groups (multiple records per patient), split by group so the same patient never appears in both train and test.

Step 4: Build a simple baseline#

Start with the simplest reasonable model โ€” predict the majority class, predict the mean, or a logistic regression. A baseline tells you whether the problem is learnable and gives a reference point. Surprisingly often, a well-tuned simple model is hard to beat on tabular data.

Step 5: Iterate#

Improve in order of expected payoff: fix data quality issues โ†’ engineer features โ†’ try stronger model families โ†’ tune hyperparameters โ†’ ensemble. Change one thing at a time and track every experiment.

Step 6: Evaluate thoroughly#

Beyond one headline metric:

  • Confusion matrix and per-class metrics.
  • Performance on important subgroups (region, gender, language) โ€” fairness checks.
  • Error analysis: read 50โ€“100 misclassified examples and categorise the causes.
  • Calibration of predicted probabilities.
  • Robustness to realistic shifts.

Step 7: Deploy and monitor#

Package the model and its preprocessing together, serve predictions, and monitor input distributions, prediction distributions and โ€” when labels arrive โ€” real-world performance. Plan for retraining. (The MLOps track covers this in depth.)

A complete example#

python
import numpy as np
import pandas as pd
from sklearn.datasets import fetch_openml
from sklearn.model_selection import train_test_split, cross_val_score
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.dummy import DummyClassifier
from sklearn.metrics import classification_report

# 1-2. Data: the Titanic survival dataset
df = fetch_openml("titanic", version=1, as_frame=True).frame
X = df[["pclass", "sex", "age", "sibsp", "parch", "fare", "embarked"]]
y = (df["survived"].astype(int))

# 3. Split first
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=0.2, stratify=y, random_state=0)

num = ["age", "sibsp", "parch", "fare"]
cat = ["pclass", "sex", "embarked"]
prep = ColumnTransformer([
    ("num", Pipeline([("impute", SimpleImputer(strategy="median")), ("scale", StandardScaler())]), num),
    ("cat", Pipeline([("impute", SimpleImputer(strategy="most_frequent")),
                      ("onehot", OneHotEncoder(handle_unknown="ignore"))]), cat),
])

# 4-5. Baseline and candidates, compared by cross-validation on the training set only
candidates = {
    "majority baseline": DummyClassifier(strategy="most_frequent"),
    "logistic regression": LogisticRegression(max_iter=1000),
    "gradient boosting": HistGradientBoostingClassifier(random_state=0),
}
for name, model in candidates.items():
    pipe = Pipeline([("prep", prep), ("model", model)])
    scores = cross_val_score(pipe, X_tr, y_tr, cv=5, scoring="f1")
    print(f"{name:<20} F1 = {scores.mean():.3f} ยฑ {scores.std():.3f}")

# 6. Final evaluation of the chosen model on the untouched test set
best = Pipeline([("prep", prep), ("model", HistGradientBoostingClassifier(random_state=0))]).fit(X_tr, y_tr)
print(classification_report(y_te, best.predict(X_te)))

Notice three professional habits: preprocessing lives inside the pipeline (so it is fitted only on training folds), we compare against a baseline, and the test set is used once.

Common workflow failures#

FailureSymptomPrevention
Data leakageSuspiciously high validation scoresSplit first; pipelines; check feature timing
Wrong metricGreat metric, useless modelFrame costs of errors up front
Test-set reuseReported results do not hold in productionValidation set for decisions; test once
Ignoring subgroupsHarm to under-represented usersSliced evaluation
No monitoringSilent degradationMonitor drift and outcomes
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Evaluation Metrics for Classification: Accuracy, Precision, Recall and F1

Accuracy can be dangerously misleading. We build the confusion matrix, define precision, recall, specificity, F-scores and balanced accuracy, and learn to choose metrics from the costs of errors.

Beginnerโฑ 5 min#061
๐Ÿ“ˆ Machine Learning

Types of Machine Learning: Supervised, Unsupervised, Self-Supervised and Reinforcement

Learning problems differ by the kind of feedback available. We map the major paradigms, their typical tasks and algorithms, and the hybrid settings โ€” semi-supervised, weak and transfer learning โ€” that dominate practice.

Beginnerโฑ 5 min#051
๐Ÿ“ˆ Machine Learning

Linear Regression from First Principles

The most important model in statistics and ML. We derive least squares geometrically and analytically, solve it with the normal equations and gradient descent, and interpret coefficients carefully.

Beginnerโฑ 5 min#053