Production MLOps Platform · Jun 2026
pipeline-lineage-tracker: Data Versioning and Lineage
Any prediction traced back through five hops to the exact raw data row
The problem
MLflow tracks experiments, DVC versions datasets and git tracks code, but none of them talk to each other. When a prediction is wrong, reconstructing its origin means piecing together separate logs and human memory.
Approach
SHA-256 content hashing gives content-addressed storage, so the same file ingested twice creates one version rather than two. A full lineage DAG lives in PostgreSQL with MinIO blob storage and NetworkX on top, where every node has a UUID and every edge is a row. Auto-tracking decorators cover pandas and sklearn, an MLflow callback links both directions, and impact analysis scores downstream risk before a data change ships.
How it works
- Every dataset is SHA-256 content-hashed before anything else, so the same file ingested twice creates one version rather than two.
- Every node carries a UUID and every edge is a row with source and target type, which keeps traversal a plain query rather than a graph database dependency.
- @track_dataframe and TrackedPipeline decorators capture pandas and sklearn work without changing the function bodies they wrap.
- An MLflow callback writes the run id into the lineage record, so lookups work in both directions.
- Impact analysis scores downstream risk from the number of affected models before a data change ships.
Key decisions
- Content addressing over version numbers
- The hash is the identity, so a model's training data can be reconstructed exactly even if the original file is long gone, and re-ingesting the same content does not create a duplicate version.
- Zero-friction instrumentation
- Decorators intercept the return value rather than asking existing pipelines to be rewritten. A team adopts this by wrapping functions, not by restructuring them, which is the difference between a tool being used and being admired.
- Answer both directions of the question
- Tracing upstream answers 'where did this prediction come from'. Impact analysis answers 'what breaks if I change this'. The second is the one that prevents incidents rather than explaining them.
Interface

Architecture
Running it
docker compose upFull setup, configuration and API reference are in the repository README.
Stack
- Python
- PostgreSQL
- MinIO
- NetworkX
- pyvis
- Streamlit
- MLflow