Skip to content

Work

Ordered by how much of the system each tier owns: reliability research, the MLOps platform underneath it, and the applied models on top.


Reliability and Evaluation Research

Production MLOps Platform

Feature store dual-store architectureData sources feed an ingestion pipeline, which writes to a PostgreSQL offline store and a Redis online store. The offline store supplies training data through a point-in-time correct join; the online store serves inference in under 10 milliseconds. A registry tracks versions across both.SourcesCSV · PG · KafkaIngestion@feature · SHA-256drift detectionOffline · PostgreSQLfull historyOnline · Redislatest value onlypoint-in-timecorrect joinTrainingno leakageServing<10msregistry · versions
The feature store: one set of definitions, two stores, and a point-in-time correct join so training data never sees the future.

Training-serving skew: features computed one way for training and another for serving. The failure is silent, it is common, and it makes offline metrics untrustworthy.

Sub-10ms online serving with point-in-time correct joins

Python · PostgreSQL · Redis · Polars · +5

Case studySourceArticle

MLflow tracks experiments, DVC versions datasets and git tracks code, but none of them talk to each other. When a prediction is wrong, reconstructing its origin means piecing together separate logs and human memory.

Any prediction traced back through five hops to the exact raw data row

Python · PostgreSQL · MinIO · NetworkX · +3

Case studySourceArticle

Applied ML and Data Science

Applied models and smaller tools. Each has a full case study.