๐Ÿ“ˆ Machine Learning ยท Lecture 35 of 47

Encoding Categorical Variables: One-Hot, Ordinal, Target and Beyond

Models need numbers, but many features are categories. We compare one-hot, ordinal, frequency, target and hashing encoders, handle high cardinality and unseen categories, and avoid target-encoding leakage.

District, occupation, language, product type, blood group โ€” tabular data is full of categorical variables. Most algorithms require numeric input, so we must encode categories as numbers. The naive approach โ€” assigning integers 1, 2, 3 โ€” can quietly damage a model. Choosing the right encoding depends on the variable's nature, its cardinality (number of distinct values) and the model.

Nominal vs ordinal#

  • Nominal categories have no order: colours, countries, languages.
  • Ordinal categories have a natural order: education level, severity (mild < moderate < severe), income bracket.

Ordinal (integer) encoding#

Map categories to integers. Appropriate for ordinal variables in the correct order.

One-hot encoding#

Create one binary column per category:

districtis_Dhakais_Chattogramis_Sylhet
Dhaka100
Sylhet001
  • The default for nominal variables with low cardinality.
  • For linear models with an intercept, drop one column (or rely on regularisation) to avoid perfect collinearity โ€” the "dummy variable trap".
  • Handle unseen categories at prediction time (handle_unknown="ignore" produces an all-zero row).
  • With high cardinality (thousands of values), one-hot creates huge sparse matrices and rare columns with little data. Group rare categories into "Other" (min_frequency in scikit-learn).

Frequency / count encoding#

Replace each category with how often it appears. Simple, one column, and often useful for trees (frequent vs rare categories behave differently). Different categories with equal counts become indistinguishable.

Target (mean) encoding#

Replace each category with the mean target for that category โ€” e.g. the default rate of each occupation. Compact and powerful for high-cardinality features, but dangerous:

  • Leakage: if a row's own label contributes to its encoding, the model sees the answer. Rare categories become near-perfect predictors on training data and fail on test data.
  • Noise: a category seen twice has an unreliable mean.

Remedies:

  1. Out-of-fold encoding โ€” compute each row's encoding from other folds only (scikit-learn's TargetEncoder does this with cross-fitting).
  2. Smoothing โ€” shrink towards the global mean:
$$ \text{enc}(c) = \frac{n_c\,\bar{y}_c + m\,\bar{y}}{n_c + m} $$

where $n_c$ is the category count and $m$ controls the strength of shrinkage โ€” a Bayesian estimate with a prior centred on the global mean.

  1. Ordered target statistics (CatBoost) โ€” encode each row using only earlier rows in a random permutation.

Hashing encoding#

Map categories to a fixed number of columns with a hash function. Memory is fixed regardless of cardinality, and new categories need no dictionary โ€” useful for streaming data and huge vocabularies. The cost is collisions (different categories sharing a column) and loss of interpretability.

Learned embeddings#

Neural networks can learn a dense vector for each category via an embedding layer, as they do for words. Similar categories (e.g. products bought together) end up close in embedding space. Entity embeddings are especially effective for high-cardinality features in deep tabular models, and the learned vectors can be reused in other models.

Comparison#

EncodingColumnsBest forRisks
Ordinal1True ordinal variables; treesFake order for nominal data
One-hot$K$Low-cardinality nominalDimensionality explosion
Frequency1Trees; high cardinalityCollisions of equal counts
Target1 (per class)High cardinalityLeakage โ€” use cross-fitting + smoothing
Hashingfixed $m$Streaming, huge cardinalityCollisions
Embedding$d$Neural networksNeeds data; less interpretable
python
import numpy as np
import pandas as pd
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder, OrdinalEncoder, TargetEncoder
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score

rng = np.random.default_rng(0)
n = 5000
occupations = [f"occ_{i}" for i in range(300)]              # high cardinality
occ_risk = dict(zip(occupations, rng.normal(0, 1, 300)))
df = pd.DataFrame({
    "education": rng.choice(["primary", "secondary", "tertiary"], n),
    "district": rng.choice(["Dhaka", "Chattogram", "Sylhet", "Khulna"], n),
    "occupation": rng.choice(occupations, n),
})
logit = df["occupation"].map(occ_risk) + (df["education"] == "tertiary") * 0.8
y = (rng.random(n) < 1 / (1 + np.exp(-logit))).astype(int)

prep = ColumnTransformer([
    ("edu", OrdinalEncoder(categories=[["primary", "secondary", "tertiary"]]), ["education"]),
    ("dist", OneHotEncoder(handle_unknown="ignore"), ["district"]),
    ("occ", TargetEncoder(target_type="binary", random_state=0), ["occupation"]),  # cross-fitted
])
pipe = make_pipeline(prep, LogisticRegression(max_iter=1000))
print("CV AUC:", cross_val_score(pipe, df, y, cv=5, scoring="roc_auc").mean().round(3))
JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST ยท Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

๐Ÿ“ˆ Machine Learning

Handling Missing Data: Mechanisms, Imputation and Indicators

Missing values are rarely random. We classify missingness as MCAR, MAR or MNAR, compare deletion and imputation strategies from simple to iterative, and show why a missingness indicator is often a feature in itself.

Intermediateโฑ 5 min#083
๐Ÿ“ˆ Machine Learning

Feature Scaling and Normalisation: Standardisation, Minโ€“Max and Robust Scaling

Many algorithms silently assume features share a scale. We explain which models need scaling and why, compare standardisation, minโ€“max, robust and quantile scaling, and show how to apply them without leakage.

Beginnerโฑ 4 min#082
๐Ÿ“ˆ Machine Learning

t-SNE and UMAP: Visualising High-Dimensional Data

Non-linear embeddings reveal cluster structure that PCA hides. We explain how t-SNE and UMAP work, what their hyperparameters do, and โ€” critically โ€” how not to misread their plots.

Intermediateโฑ 5 min#079