Skip to content

About


Most ML work stops at the model. I spend most of mine on what surrounds it: the evaluation harnesses, feature stores, lineage trackers and rollback systems that decide whether something which worked in a notebook keeps working for users. That came from measuring things other people assume. Whether structured output enforcement actually holds across chained calls. Whether a cited source really says what the answer claims. The answer is often no, and the corrections that came out of finding that are on this site rather than quietly fixed.

  • Production ML fails at the seams rather than the model: training-serving skew, silent drift, predictions that can't be traced to the data that caused them.
  • A model that can't be measured can't be trusted, so evaluation harnesses come before accuracy numbers.
  • Infrastructure isn't a tool until someone else adopts it, which is why these ship to PyPI, npm and Homebrew rather than living on a branch.

I came to this from biochemistry, which is a longer route than most people take, and probably why I care as much as I do about whether a measurement actually holds up. Away from work I play far too much football on a console, and support Manchester United, which is its own long lesson in evaluating performance honestly.

Portrait of Emmanuel Nwanguma

Experience

  1. Independent ML/MLOps Engineer

    Emart AI

    Jul 2024 – Present · Remote, Lagos

    • Founded and maintain a six-tool open-source AI infrastructure project spanning agent memory, evaluation, cost observability and credential scoping.
    • Built a production MLOps platform covering feature store, pipeline lineage tracking and canary deployment with automated rollback.
    • Published technical articles in Towards AI, DataDrivenInvestor and NextGenAI, documenting production ML failure modes and their fixes.
  2. Open Source Contributor

    aden-hive/hive (OpenHive)

    Jan 2026 – Present · Remote

    • #13 of 200+ contributors on an 11,000-star multi-agent framework, with 14 merged pull requests.
    • Added 458 tests across 7 security scanning tools and 3 API integration tools, closing coverage gaps flagged by maintainers.
    • Shipped a BigQuery MCP tool integration, fixed credential exception handling and an MCP resource leak, and wrote 33 tool READMEs.
  3. Health Data Specialist

    Mido Health Diagnostic

    Nov 2022 – Mar 2025 · Abia State, Nigeria

    • Owned data quality for clinical datasets of 50,000+ records, building automated Python and SQL validation pipelines that held accuracy above 99%.
    • Cut manual corrections by roughly 30% while maintaining strict healthcare data privacy compliance.
    • Worked directly with clinicians, lab technicians and administrators to resolve data quality issues at source.
  4. Data Analyst Intern

    ACME Software Lab

    Jan 2024 – Apr 2024 · Remote

    • Built a reusable ETL pipeline for cleaning, transformation and validation across multiple sources, adopted as the team's standard template.
    • Applied statistical analysis to translate raw data into insights used in stakeholder decisions.
  5. Data & Records Assistant

    Ministry of Health, Osogbo — HIV/AIDS Unit (NYSC)

    Oct 2021 – Oct 2022 · Osogbo, Nigeria

    • Managed records and monthly statistical reporting for 150+ HIV/AIDS patients using Excel and SPSS, informing public health planning.
    • Redesigned filing systems, improving record retrieval efficiency by 40% while maintaining strict patient confidentiality.

Education

  • MSc Financial Engineering

    In progress

    WorldQuant University

    Resuming Oct 2026

  • BSc Biochemistry — Second Class Upper

    Clifford University, Abia State

    Nov 2016 – Mar 2021

Publication

A Systematic Review of Generative AI in Education

Journal of Computer Sciences and Applications, 2024, 12(1), 25–30 · DOI 10.12691/jcsa-12-1-4 · Open access · co-authored

Certifications

Stack

Languages
Python · Go · TypeScript · SQL · Bash
Methods
Experiment design · A/B testing (frequentist and Bayesian) · Causal inference · Feature engineering · Statistical testing · Error analysis · LLM evaluation and regression testing · Drift detection
ML and AI
PyTorch · Hugging Face Transformers · PEFT/LoRA/QLoRA · scikit-learn · FAISS · ChromaDB · LangChain · LlamaIndex · Evidently AI · MCP
LLM APIs
OpenAI · Anthropic · Groq · Gemini · Ollama
MLOps and Deployment
Docker · MLflow · FastAPI · GitHub Actions · Prometheus · Grafana · Bicep · Azure ML · Microsoft Foundry
Cloud and Data
AWS (SageMaker, EC2, S3) · GCP BigQuery · Modal · Vercel · PostgreSQL · Redis · Kafka/Redpanda · TimescaleDB · MinIO