📈 Machine Learning · Lecture 32 of 47

Feature Engineering: Turning Raw Data into Signal

Better features beat better algorithms. We survey the craft — transformations, interactions, aggregations, date and text features, domain-driven ratios — and how to engineer features without leaking the target.

Andrew Ng once remarked that applied machine learning is basically feature engineering. Deep learning has automated feature learning for images, audio and text, but for tabular data — the most common data in business, government and humanitarian work — thoughtful features remain the single biggest lever on model quality. Today we study the craft.

What makes a good feature?#

A good feature is:

  • Informative — related to the target.
  • Available at prediction time — no leakage from the future or the label.
  • Robust — stable over time and across populations.
  • Simple to compute and explain where possible.

1. Numeric transformations#

  • Log / power transforms for skewed variables (income, population, counts): $\log(1 + x)$ or Box–Cox/Yeo–Johnson. They stabilise variance and make relationships more linear for linear models.
  • Binning into quantiles or domain-meaningful ranges (age groups) — can capture non-linearity for linear models and improve robustness to outliers, at the cost of information.
  • Clipping (winsorising) extreme values.
  • Ratios and differences often carry more meaning than raw values: debt-to-income, price per square metre, household size per room, change since last month.

2. Interaction features#

Products and combinations of features capture effects that depend on context: rainfall × temperature for crop yield, is_weekend × hour for demand. Linear models cannot discover interactions on their own; trees and neural networks can, but explicit interactions still help them learn faster with less data.

3. Date and time features#

A timestamp is not a useful raw number. Extract: hour, day of week, month, quarter, holiday flags, days until/since an event, season, and cyclical encodings so that December is next to January:

$$ \text{month}_{\sin} = \sin\left(\frac{2\pi\,\text{month}}{12}\right), \qquad \text{month}_{\cos} = \cos\left(\frac{2\pi\,\text{month}}{12}\right) $$

4. Aggregation (group) features#

For relational data — transactions per customer, visits per patient, registrations per location — compute statistics over related records: counts, sums, means, maxima, standard deviations, recency (time since last event), frequency and trends. Aggregations over time windows (last 7, 30, 90 days) are extremely powerful.

5. Text features#

Word counts, TF-IDF vectors, text length, presence of keywords, sentiment scores, and — increasingly — embeddings from pretrained language models, which convert free text into dense vectors that capture meaning.

6. Geographic features#

Distances to key locations (nearest clinic, market, water point), density of points within a radius, administrative-region encodings, elevation, and features derived from satellite imagery.

7. Domain knowledge features#

The most valuable features usually come from talking to domain experts. A nutritionist knows that the ratio of weight to height predicts malnutrition better than either alone (that is why weight-for-height z-scores exist). A logistics officer knows that road access in the rainy season matters more than distance. No algorithm will reliably discover what an expert already knows from a few thousand examples.

A worked example#

python
import numpy as np
import pandas as pd

rng = np.random.default_rng(0)
n = 1000
df = pd.DataFrame({
    "household_id": rng.integers(0, 200, n),
    "visit_time": pd.to_datetime("2025-01-01") + pd.to_timedelta(rng.integers(0, 365 * 24, n), unit="h"),
    "members": rng.integers(1, 10, n),
    "rooms": rng.integers(1, 5, n),
    "monthly_income": rng.lognormal(9, 0.8, n),
})
df = df.sort_values("visit_time")

# Numeric transforms and ratios
df["log_income"] = np.log1p(df["monthly_income"])
df["income_per_member"] = df["monthly_income"] / df["members"]
df["crowding"] = df["members"] / df["rooms"]

# Date features with cyclical encoding
df["hour"] = df["visit_time"].dt.hour
df["dow"] = df["visit_time"].dt.dayofweek
df["month_sin"] = np.sin(2 * np.pi * df["visit_time"].dt.month / 12)
df["month_cos"] = np.cos(2 * np.pi * df["visit_time"].dt.month / 12)

# Point-in-time aggregations: only PAST visits of the same household
g = df.groupby("household_id")
df["prior_visits"] = g.cumcount()
df["days_since_last_visit"] = g["visit_time"].diff().dt.total_seconds() / 86400
df["prior_mean_income"] = g["monthly_income"].transform(lambda s: s.shift().expanding().mean())

print(df.head(8).round(2).to_string())

Note the shift() in the aggregation: each row sees only earlier visits. Without it, the feature would include the current (and future) visits — leakage.

Automated feature engineering#

Tools like Featuretools (deep feature synthesis) generate many aggregation features automatically from relational tables. They are useful for exploration, but they generate hundreds of candidates — combine them with feature selection and careful leakage checks.

A disciplined process#

  1. Start with a baseline on raw features.
  2. Brainstorm features with domain experts; write down the hypothesis behind each.
  3. Add features in groups; measure validation improvement for each group.
  4. Check importance and error analysis to inspire the next features.
  5. Keep all feature code in a reproducible pipeline shared by training and serving, so the model sees identical features in production.
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

📈 Machine Learning

Polynomial Regression and Basis Functions: Non-Linearity with Linear Models

Linear models can fit curves if we transform the inputs. We study polynomial, spline and radial basis features, watch overfitting happen as degree grows, and connect it to model selection.

Beginner⏱ 5 min#054
📈 Machine Learning

Linear Discriminant Analysis: Supervised Dimensionality Reduction

Unlike PCA, LDA uses labels to find projections that separate classes. We derive Fisher's criterion, the generative Gaussian view, and compare LDA with PCA, QDA and logistic regression.

Intermediate⏱ 4 min#080
📈 Machine Learning

Feature Scaling and Normalisation: Standardisation, Min–Max and Robust Scaling

Many algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, min–max, robust and quantile scaling, and show how to apply them without leakage.

Beginner⏱ 4 min#082