Andrew Ng once remarked that applied machine learning is basically feature engineering. Deep learning has automated feature learning for images, audio and text, but for tabular data — the most common data in business, government and humanitarian work — thoughtful features remain the single biggest lever on model quality. Today we study the craft.
What makes a good feature?#
A good feature is:
- Informative — related to the target.
- Available at prediction time — no leakage from the future or the label.
- Robust — stable over time and across populations.
- Simple to compute and explain where possible.
1. Numeric transformations#
- Log / power transforms for skewed variables (income, population, counts): $\log(1 + x)$ or Box–Cox/Yeo–Johnson. They stabilise variance and make relationships more linear for linear models.
- Binning into quantiles or domain-meaningful ranges (age groups) — can capture non-linearity for linear models and improve robustness to outliers, at the cost of information.
- Clipping (winsorising) extreme values.
- Ratios and differences often carry more meaning than raw values: debt-to-income, price per square metre, household size per room, change since last month.
2. Interaction features#
Products and combinations of features capture effects that depend on context: rainfall × temperature for crop yield, is_weekend × hour for demand. Linear models cannot discover interactions on their own; trees and neural networks can, but explicit interactions still help them learn faster with less data.
3. Date and time features#
A timestamp is not a useful raw number. Extract: hour, day of week, month, quarter, holiday flags, days until/since an event, season, and cyclical encodings so that December is next to January:
4. Aggregation (group) features#
For relational data — transactions per customer, visits per patient, registrations per location — compute statistics over related records: counts, sums, means, maxima, standard deviations, recency (time since last event), frequency and trends. Aggregations over time windows (last 7, 30, 90 days) are extremely powerful.
5. Text features#
Word counts, TF-IDF vectors, text length, presence of keywords, sentiment scores, and — increasingly — embeddings from pretrained language models, which convert free text into dense vectors that capture meaning.
6. Geographic features#
Distances to key locations (nearest clinic, market, water point), density of points within a radius, administrative-region encodings, elevation, and features derived from satellite imagery.
7. Domain knowledge features#
The most valuable features usually come from talking to domain experts. A nutritionist knows that the ratio of weight to height predicts malnutrition better than either alone (that is why weight-for-height z-scores exist). A logistics officer knows that road access in the rainy season matters more than distance. No algorithm will reliably discover what an expert already knows from a few thousand examples.
A worked example#
import numpy as np
import pandas as pd
rng = np.random.default_rng(0)
n = 1000
df = pd.DataFrame({
"household_id": rng.integers(0, 200, n),
"visit_time": pd.to_datetime("2025-01-01") + pd.to_timedelta(rng.integers(0, 365 * 24, n), unit="h"),
"members": rng.integers(1, 10, n),
"rooms": rng.integers(1, 5, n),
"monthly_income": rng.lognormal(9, 0.8, n),
})
df = df.sort_values("visit_time")
# Numeric transforms and ratios
df["log_income"] = np.log1p(df["monthly_income"])
df["income_per_member"] = df["monthly_income"] / df["members"]
df["crowding"] = df["members"] / df["rooms"]
# Date features with cyclical encoding
df["hour"] = df["visit_time"].dt.hour
df["dow"] = df["visit_time"].dt.dayofweek
df["month_sin"] = np.sin(2 * np.pi * df["visit_time"].dt.month / 12)
df["month_cos"] = np.cos(2 * np.pi * df["visit_time"].dt.month / 12)
# Point-in-time aggregations: only PAST visits of the same household
g = df.groupby("household_id")
df["prior_visits"] = g.cumcount()
df["days_since_last_visit"] = g["visit_time"].diff().dt.total_seconds() / 86400
df["prior_mean_income"] = g["monthly_income"].transform(lambda s: s.shift().expanding().mean())
print(df.head(8).round(2).to_string())Note the shift() in the aggregation: each row sees only earlier visits. Without it, the feature would include the current (and future) visits — leakage.
Automated feature engineering#
Tools like Featuretools (deep feature synthesis) generate many aggregation features automatically from relational tables. They are useful for exploration, but they generate hundreds of candidates — combine them with feature selection and careful leakage checks.
A disciplined process#
- Start with a baseline on raw features.
- Brainstorm features with domain experts; write down the hypothesis behind each.
- Add features in groups; measure validation improvement for each group.
- Check importance and error analysis to inspire the next features.
- Keep all feature code in a reproducible pipeline shared by training and serving, so the model sees identical features in production.