Skip to content

What Broke

Every project on this site shipped a measurement. These are the bugs that would have made those measurements wrong, caught before publication and documented because a result you cannot audit is not a result.


so-agent

  • An earlier version of Customer.name was a required non-nullable string. Eight of the ten test tickets name nobody, so models had no legal way to say "not stated", and the benchmark recorded 80% of extractions as inventing a customer name. They were not inventing: 292 of the flagged values were the literal string the field description had asked for. With the field nullable and the description asking for null, grounding is 100% on every model except the 7B one. Give a model a way to decline and it takes it.
  • The LLM critic scoring semantic accuracy originally agreed with hand labels only 50% of the time, which is chance. It could not see the allowed enum values and treated inferred judgements like priority as facts to be found in the text. That version would have published a 91.7% failure rate for a system whose real rate was 41.7%. An unmeasured judge is just a second unvalidated model.

ragwell

  • A substring match on money, where N5,000,000 contains 500,000, scored every correct answer to the central question as wrong. It would have inverted the headline.
  • A partial benchmark run suggested the opposite conclusion on superseded documents and was nearly published.

personalization-feedback-engine

  • A collaborative filtering model scored below the random floor because one coincidental co-click created a similarity.

causal-inference-engine

  • The covariate classifier labelled an instrument as a confounder. The logic asked whether the variable reaches the outcome, but every cause of the treatment reaches the outcome through the treatment. The fix was to ask whether a path exists that avoids the treatment. Small distinction, completely different adjustment set.
  • The validation harness graded the instrumental row as a failure, when that row exists precisely to show adjustment has limits. Reporting a correct demonstration as a defect is wrong. Rows now declare their own intent: recover, improve, or fail.
  • The placebo test compared the placebo estimate against a fraction of the original, which fails every correct null result since that fraction is near zero by definition. It reported a correctly detected absence of effect as a failure. The fix was to judge against the spread of placebo estimates themselves.

ml-feature-store

  • A SQLAlchemy async engine created at module import with a connection pool, then used inside a Uvicorn lifespan handler, raised RuntimeError: Task got Future attached to a different loop. The pool binds to whichever event loop touches it first, and Uvicorn creates its own. The fix was NullPool.
  • Point-in-time correctness looked simple and wasn't. No feature value before the label timestamp, multiple values on the same timestamp, timezone handling: each cost a debugging session.

ml-canary-deploy

  • The canary traffic split was read once at startup and never again, so a canary started later would have received zero traffic forever, indistinguishable from a healthy deployment.

realtime-anomaly-detection

  • Simulated data returned an implausible ~1.00 precision and recall. The pipeline was revalidated against a real Kaggle fraud dataset, and the simulation result was published as a warning rather than as a result.

Every entry here is a defect found in my own work before the result was published, and each one changed what got reported. Projects with nothing worth reporting are simply absent.