Imagine a feature "number of assistance requests by this household in the last 90 days". The data scientist computes it in a notebook with SQL for training. Months later, an engineer re-implements it in the serving system — with a subtly different time window, time zone or treatment of cancelled requests. The model now receives different inputs in production than in training, and its accuracy quietly drops. This training–serving skew is one of the most common and damaging bugs in production ML. Feature stores exist largely to prevent it.
What is a feature store?#
A feature store is a data system that manages ML features as reusable, documented, versioned assets, with:
- Feature definitions written once (transformations, entity keys, time windows).
- An offline store holding historical feature values for training — large, append-only, often in a data warehouse or lakehouse.
- An online store holding the latest feature values for low-latency serving — a key–value database (e.g. Redis, DynamoDB, Bigtable).
- Materialisation jobs that compute features and keep both stores in sync.
- Point-in-time correct retrieval for building training sets.
- A registry with metadata, ownership, documentation and lineage.
Examples: Feast (open source), Hopsworks, Tecton, and feature stores in cloud ML platforms (Vertex AI, SageMaker, Databricks).
Point-in-time correctness#
When building a training set of labelled events (e.g. "was this household later found to be in urgent need?" at time $t$), each row must use feature values as they were known at time $t$ — not later values. Otherwise future information leaks into training, inflating offline metrics that the production model cannot reproduce.
A naive join of the event table with the latest feature table is wrong. A correct "as-of" join selects, for each event, the most recent feature value with timestamp $\le$ the event time:
import pandas as pd
events = pd.DataFrame({
"household_id": ["H1", "H1", "H2"],
"event_time": pd.to_datetime(["2025-03-01", "2025-06-01", "2025-04-15"]),
"label": [0, 1, 0]})
features = pd.DataFrame({
"household_id": ["H1", "H1", "H1", "H2"],
"feature_time": pd.to_datetime(["2025-02-01", "2025-05-01", "2025-07-01", "2025-04-01"]),
"requests_90d": [1, 4, 9, 2]})
training = pd.merge_asof(events.sort_values("event_time"), features.sort_values("feature_time"),
left_on="event_time", right_on="feature_time", by="household_id",
direction="backward") # only values known at event time
print(training[["household_id", "event_time", "requests_90d", "label"]])
# H1 on 2025-06-01 gets 4 (from May), NOT 9 (from July, which would be leakage)Feature stores implement this at scale, including features with their own processing delays (a value computed on the 1st may only be available on the 3rd).
A Feast example#
# feature_repo/features.py
from datetime import timedelta
from feast import Entity, FeatureView, Field, FileSource
from feast.types import Int64, Float32
household = Entity(name="household", join_keys=["household_id"])
source = FileSource(path="data/household_features.parquet", timestamp_field="feature_time")
household_stats = FeatureView(
name="household_stats",
entities=[household],
ttl=timedelta(days=120),
schema=[Field(name="requests_90d", dtype=Int64), Field(name="avg_income_6m", dtype=Float32)],
source=source,
)from feast import FeatureStore
store = FeatureStore(repo_path="feature_repo")
# Offline: point-in-time correct training data
train_df = store.get_historical_features(entity_df=events.rename(columns={"event_time": "event_timestamp"}),
features=["household_stats:requests_90d",
"household_stats:avg_income_6m"]).to_df()
# Online: latest values at prediction time (after `feast materialize`)
online = store.get_online_features(features=["household_stats:requests_90d"],
entity_rows=[{"household_id": "H1"}]).to_dict()The same definitions feed both training and serving — eliminating skew by construction.
Benefits#
- Consistency between training and serving.
- Reuse across teams and models; less duplicated feature engineering.
- Discoverability: a catalogue of documented features with owners.
- Governance: access control on sensitive features; lineage for audits.
- Low-latency serving of precomputed aggregates.
When is a feature store worth it?#
Feature stores add infrastructure and operational cost. They pay off when you have many models sharing features, real-time serving needs with complex aggregates, multiple teams, and recurring skew or leakage bugs. A single batch model can often achieve consistency more simply: a shared Python feature module used by both training and batch scoring, plus point-in-time joins in SQL.