⚙️ MLOps & Engineering · Lecture 9 of 15

Feature Stores and Training–Serving Consistency

Feature stores manage features as shared, versioned assets available offline for training and online for serving. We explain training–serving skew, point-in-time correctness, online vs offline stores, and when a feature store is worth it.

Imagine a feature "number of assistance requests by this household in the last 90 days". The data scientist computes it in a notebook with SQL for training. Months later, an engineer re-implements it in the serving system — with a subtly different time window, time zone or treatment of cancelled requests. The model now receives different inputs in production than in training, and its accuracy quietly drops. This training–serving skew is one of the most common and damaging bugs in production ML. Feature stores exist largely to prevent it.

What is a feature store?#

A feature store is a data system that manages ML features as reusable, documented, versioned assets, with:

  1. Feature definitions written once (transformations, entity keys, time windows).
  2. An offline store holding historical feature values for training — large, append-only, often in a data warehouse or lakehouse.
  3. An online store holding the latest feature values for low-latency serving — a key–value database (e.g. Redis, DynamoDB, Bigtable).
  4. Materialisation jobs that compute features and keep both stores in sync.
  5. Point-in-time correct retrieval for building training sets.
  6. A registry with metadata, ownership, documentation and lineage.

Examples: Feast (open source), Hopsworks, Tecton, and feature stores in cloud ML platforms (Vertex AI, SageMaker, Databricks).

Point-in-time correctness#

When building a training set of labelled events (e.g. "was this household later found to be in urgent need?" at time $t$), each row must use feature values as they were known at time $t$ — not later values. Otherwise future information leaks into training, inflating offline metrics that the production model cannot reproduce.

A naive join of the event table with the latest feature table is wrong. A correct "as-of" join selects, for each event, the most recent feature value with timestamp $\le$ the event time:

python
import pandas as pd

events = pd.DataFrame({
    "household_id": ["H1", "H1", "H2"],
    "event_time": pd.to_datetime(["2025-03-01", "2025-06-01", "2025-04-15"]),
    "label": [0, 1, 0]})
features = pd.DataFrame({
    "household_id": ["H1", "H1", "H1", "H2"],
    "feature_time": pd.to_datetime(["2025-02-01", "2025-05-01", "2025-07-01", "2025-04-01"]),
    "requests_90d": [1, 4, 9, 2]})

training = pd.merge_asof(events.sort_values("event_time"), features.sort_values("feature_time"),
                         left_on="event_time", right_on="feature_time", by="household_id",
                         direction="backward")                 # only values known at event time
print(training[["household_id", "event_time", "requests_90d", "label"]])
# H1 on 2025-06-01 gets 4 (from May), NOT 9 (from July, which would be leakage)

Feature stores implement this at scale, including features with their own processing delays (a value computed on the 1st may only be available on the 3rd).

A Feast example#

python
# feature_repo/features.py
from datetime import timedelta
from feast import Entity, FeatureView, Field, FileSource
from feast.types import Int64, Float32

household = Entity(name="household", join_keys=["household_id"])
source = FileSource(path="data/household_features.parquet", timestamp_field="feature_time")

household_stats = FeatureView(
    name="household_stats",
    entities=[household],
    ttl=timedelta(days=120),
    schema=[Field(name="requests_90d", dtype=Int64), Field(name="avg_income_6m", dtype=Float32)],
    source=source,
)
python
from feast import FeatureStore
store = FeatureStore(repo_path="feature_repo")
# Offline: point-in-time correct training data
train_df = store.get_historical_features(entity_df=events.rename(columns={"event_time": "event_timestamp"}),
                                         features=["household_stats:requests_90d",
                                                   "household_stats:avg_income_6m"]).to_df()
# Online: latest values at prediction time (after `feast materialize`)
online = store.get_online_features(features=["household_stats:requests_90d"],
                                   entity_rows=[{"household_id": "H1"}]).to_dict()

The same definitions feed both training and serving — eliminating skew by construction.

Benefits#

  • Consistency between training and serving.
  • Reuse across teams and models; less duplicated feature engineering.
  • Discoverability: a catalogue of documented features with owners.
  • Governance: access control on sensitive features; lineage for audits.
  • Low-latency serving of precomputed aggregates.

When is a feature store worth it?#

Feature stores add infrastructure and operational cost. They pay off when you have many models sharing features, real-time serving needs with complex aggregates, multiple teams, and recurring skew or leakage bugs. A single batch model can often achieve consistency more simply: a shared Python feature module used by both training and batch scoring, plus point-in-time joins in SQL.

JA
Written by

Janin A Apurba

B.Sc. in CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. Teaching AI, ML and Deep Learning to the next generation of engineers and researchers.

Keep learning

Related lectures

⚙️ MLOps & Engineering

CI/CD for Machine Learning: Testing and Automating ML Systems

Continuous integration and delivery bring software-engineering discipline to ML. We cover the testing pyramid for ML — code, data and model tests — quality gates, continuous training, and a practical GitHub Actions workflow.

Intermediate⏱ 5 min#249
⚙️ MLOps & Engineering

Edge AI and TinyML: Running Models on Phones and Microcontrollers

Running models on-device brings privacy, offline operation, low latency and low cost. We cover the edge hardware spectrum, the optimisation pipeline, TensorFlow Lite, ONNX Runtime and TinyML on microcontrollers, and field-deployment lessons.

Intermediate⏱ 5 min#251
⚙️ MLOps & Engineering

Model Monitoring: Detecting Data Drift and Concept Drift

Deployed models degrade silently as the world changes. We define data drift, concept drift and prediction drift, measure drift with PSI and statistical tests, monitor without labels, and design alerts and retraining triggers.

Intermediate⏱ 5 min#248