Skip to content

Emmanuel Nwanguma · Lagos, Nigeria · Open to remote worldwide

I build ML systems and the evidence they actually work.

Feature stores, lineage tracking, evaluation harnesses, canary rollback, and the models they carry. Six open-source tools shipped, and every result published with the run that broke it.

Data Science · ML Engineering · MLOps · AI Engineering

Portrait of Emmanuel Nwanguma
6
open-source tools shipped · 5 on PyPI, 1 on npm
458
tests added to aden-hive/hive, #13 of 200+ contributors
AI-300
Microsoft Certified ML Operations Engineer Associate

01 / Selected work

Training-serving skew: features computed one way for training and another for serving. The failure is silent, it is common, and it makes offline metrics untrustworthy.

Sub-10ms online serving with point-in-time correct joins

Python · PostgreSQL · Redis · Polars · +5

Case studySourceArticle
All 15 projects

02 / What broke

Every project here shipped a measurement. These are the bugs that would have made those measurements wrong, caught before publication.

so-agent

An earlier version of Customer.name was a required non-nullable string. Eight of the ten test tickets name nobody, so models had no legal way to say "not stated", and the benchmark recorded 80% of extractions as inventing a customer name. They were not inventing: 292 of the flagged values were the literal string the field description had asked for. With the field nullable and the description asking for null, grounding is 100% on every model except the 7B one. Give a model a way to decline and it takes it.

ragwell

A substring match on money, where N5,000,000 contains 500,000, scored every correct answer to the central question as wrong. It would have inverted the headline.

All 12 corrections

03 / Latest writing