Skip to content

Production MLOps Platform · Jun 2026

pipeline-lineage-tracker: Data Versioning and Lineage

Any prediction traced back through five hops to the exact raw data row

Source

The problem

MLflow tracks experiments, DVC versions datasets and git tracks code, but none of them talk to each other. When a prediction is wrong, reconstructing its origin means piecing together separate logs and human memory.

Approach

SHA-256 content hashing gives content-addressed storage, so the same file ingested twice creates one version rather than two. A full lineage DAG lives in PostgreSQL with MinIO blob storage and NetworkX on top, where every node has a UUID and every edge is a row. Auto-tracking decorators cover pandas and sklearn, an MLflow callback links both directions, and impact analysis scores downstream risk before a data change ships.

How it works

Key decisions

Content addressing over version numbers
The hash is the identity, so a model's training data can be reconstructed exactly even if the original file is long gone, and re-ingesting the same content does not create a duplicate version.
Zero-friction instrumentation
Decorators intercept the return value rather than asking existing pipelines to be rewritten. A team adopts this by wrapping functions, not by restructuring them, which is the difference between a tool being used and being admired.
Answer both directions of the question
Tracing upstream answers 'where did this prediction come from'. Impact analysis answers 'what breaks if I change this'. The second is the one that prevents incidents rather than explaining them.

Interface

The interactive lineage DAG, colour-coded by node type
Ten nodes and nine edges across four node types. Datasets, pipeline runs, models and predictions in one graph, traceable in either direction.

Architecture

Pipeline lineage DAGFive hops connect a prediction back to its raw data: raw data, ingestion, features, training run, model, prediction. Every artifact is content-hashed with SHA-256 so the chain can be traversed in either direction.Raw dataCSV rowIngestionSHA-256FeaturesversionedTraininggit commitPredictionloggedtrace upstream — 5 hops, one command

Running it

docker compose up

Full setup, configuration and API reference are in the repository README.

Stack

Written up in