Production MLOps Platform · May – Jun 2026
ml-feature-store: Lightweight Feature Store
Sub-10ms online serving with point-in-time correct joins
The problem
Training-serving skew: features computed one way for training and another for serving. The failure is silent, it is common, and it makes offline metrics untrustworthy.
Approach
A dual-store architecture (PostgreSQL offline, Redis online) with LEFT JOIN LATERAL point-in-time correct joins, so training data reflects only what was knowable at the label timestamp. A @feature decorator SHA-256 hashes the function body, so changing a weighting automatically creates a new version with no manual version bumps. Schema validation and Evidently AI drift detection run on every ingestion run. Ships with a CLI, a Python SDK and a Streamlit registry.
How it works
- A @feature decorator SHA-256 hashes the function body, so changing a weighting creates a new version automatically, with no manual version bumps to forget.
- Decorated functions register themselves, so ingestion discovers features by scanning the registry rather than a hand-maintained list.
- Point-in-time joins use LEFT JOIN LATERAL with computed_at <= the label timestamp, returning only the most recent value knowable at that moment.
- A scheduled job syncs values from the PostgreSQL offline store to the Redis online store, full or incremental or per feature.
- Evidently AI compares every ingestion run against the reference distribution and stores the drift score with the run.
Key decisions
- One definition, two backing stores
- Training and serving read the same feature definitions from stores optimised for different access patterns: full history in PostgreSQL, latest value in Redis. That shared definition is what actually removes training-serving skew.
- Hash the code, not the output
- Versioning the function body catches a changed weighting even when the resulting values look similar, and removes the class of bug where someone forgets to bump a version number.
- NullPool instead of a shared connection pool
- A SQLAlchemy async engine created at import time binds to whichever event loop touches it first, and Uvicorn creates its own. Each request now opens and closes its own connection: slightly slower, no cross-loop state.
Architecture
What broke
- A SQLAlchemy async engine created at module import with a connection pool, then used inside a Uvicorn lifespan handler, raised RuntimeError: Task got Future attached to a different loop. The pool binds to whichever event loop touches it first, and Uvicorn creates its own. The fix was NullPool.
- Point-in-time correctness looked simple and wasn't. No feature value before the label timestamp, multiple values on the same timestamp, timezone handling: each cost a debugging session.
Running it
docker compose upFull setup, configuration and API reference are in the repository README.
Stack
- Python
- PostgreSQL
- Redis
- Polars
- FastAPI
- Click
- APScheduler
- Evidently AI
- Streamlit