District, occupation, language, product type, blood group โ tabular data is full of categorical variables. Most algorithms require numeric input, so we must encode categories as numbers. The naive approach โ assigning integers 1, 2, 3 โ can quietly damage a model. Choosing the right encoding depends on the variable's nature, its cardinality (number of distinct values) and the model.
Nominal vs ordinal#
- Nominal categories have no order: colours, countries, languages.
- Ordinal categories have a natural order: education level, severity (mild < moderate < severe), income bracket.
Ordinal (integer) encoding#
Map categories to integers. Appropriate for ordinal variables in the correct order.
One-hot encoding#
Create one binary column per category:
| district | is_Dhaka | is_Chattogram | is_Sylhet |
|---|---|---|---|
| Dhaka | 1 | 0 | 0 |
| Sylhet | 0 | 0 | 1 |
- The default for nominal variables with low cardinality.
- For linear models with an intercept, drop one column (or rely on regularisation) to avoid perfect collinearity โ the "dummy variable trap".
- Handle unseen categories at prediction time (
handle_unknown="ignore"produces an all-zero row). - With high cardinality (thousands of values), one-hot creates huge sparse matrices and rare columns with little data. Group rare categories into "Other" (
min_frequencyin scikit-learn).
Frequency / count encoding#
Replace each category with how often it appears. Simple, one column, and often useful for trees (frequent vs rare categories behave differently). Different categories with equal counts become indistinguishable.
Target (mean) encoding#
Replace each category with the mean target for that category โ e.g. the default rate of each occupation. Compact and powerful for high-cardinality features, but dangerous:
- Leakage: if a row's own label contributes to its encoding, the model sees the answer. Rare categories become near-perfect predictors on training data and fail on test data.
- Noise: a category seen twice has an unreliable mean.
Remedies:
- Out-of-fold encoding โ compute each row's encoding from other folds only (scikit-learn's
TargetEncoderdoes this with cross-fitting). - Smoothing โ shrink towards the global mean:
where $n_c$ is the category count and $m$ controls the strength of shrinkage โ a Bayesian estimate with a prior centred on the global mean.
- Ordered target statistics (CatBoost) โ encode each row using only earlier rows in a random permutation.
Hashing encoding#
Map categories to a fixed number of columns with a hash function. Memory is fixed regardless of cardinality, and new categories need no dictionary โ useful for streaming data and huge vocabularies. The cost is collisions (different categories sharing a column) and loss of interpretability.
Learned embeddings#
Neural networks can learn a dense vector for each category via an embedding layer, as they do for words. Similar categories (e.g. products bought together) end up close in embedding space. Entity embeddings are especially effective for high-cardinality features in deep tabular models, and the learned vectors can be reused in other models.
Comparison#
| Encoding | Columns | Best for | Risks |
|---|---|---|---|
| Ordinal | 1 | True ordinal variables; trees | Fake order for nominal data |
| One-hot | $K$ | Low-cardinality nominal | Dimensionality explosion |
| Frequency | 1 | Trees; high cardinality | Collisions of equal counts |
| Target | 1 (per class) | High cardinality | Leakage โ use cross-fitting + smoothing |
| Hashing | fixed $m$ | Streaming, huge cardinality | Collisions |
| Embedding | $d$ | Neural networks | Needs data; less interpretable |
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(0)
n = 5000
occupations = [f"occ_{i}" for i in range(300)] # high cardinality
occ_risk = dict(zip(occupations, rng.normal(0, 1, 300)))
df = pd.DataFrame({
"education": rng.choice(["primary", "secondary", "tertiary"], n),
"district": rng.choice(["Dhaka", "Chattogram", "Sylhet", "Khulna"], n),
"occupation": rng.choice(occupations, n),
})
logit = df["occupation"].map(occ_risk) + (df["education"] == "tertiary") * 0.8
y = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int)
prep = ColumnTransformer([
("edu", OrdinalEncoder(categories=[["primary", "secondary", "tertiary"]]), ["education"]),
("dist", OneHotEncoder(handle_unknown="ignore"), ["district"]),
("occ", TargetEncoder(target_type="binary", random_state=0), ["occupation"]), # cross-fitted
])
pipe = make_pipeline(prep, LogisticRegression(max_iter=1000))
print("CV AUC:", cross_val_score(pipe, df, y, cv=5, scoring="roc_auc").mean().round(3))