Reliability and Evaluation Research · Aug 2026
so-agent: Structured Output Reliability Benchmark
87.6% single-call success collapses to 26.7% over ten chained calls
The problem
Everyone assumes structured output enforcement works. Nobody measures whether providers actually honour it.
Approach
A capability probe distinguishing four outcomes per enforcement tier. The critical one is IGNORED, where a provider accepts a directive and silently does not apply it. A three-tier enforcement ladder selected from the probe. A typed failure taxonomy separating free repairs from unretryable refusals. Every rate published with its confidence interval and n, across 833 tests on 6 models, backed by a 340-test suite.
How it works
- A capability probe makes a real API call per model per enforcement tier and records four outcomes. The critical one is IGNORED, where a provider accepts the directive and silently does not apply it.
- Probe results are committed as a dated capabilities.json, because provider behaviour moves. Two models measured on 14 August no longer existed on 18 August.
- A three-tier enforcement ladder (json_schema, json_object, prompt_only) selects the strongest tier the probe actually confirmed for that model rather than the strongest one documented.
- Structural validity and semantic accuracy are scored separately, against hand-written labels that permit a set of answers wherever a ticket genuinely straddles two teams.
- Every rate is published with its confidence interval and n, across 833 tests on 6 models.
Key decisions
- Probe the provider, do not trust the documentation
- Every cell in the capability matrix is a real API call. Three of six models on one provider reject a native JSON schema outright, which is why the ladder exists rather than being a formality: on half the line-up there is nothing to fall back from.
- Measure structure and meaning separately
- gpt-oss-20b produces structurally valid output 100% of the time at every tier and is semantically accurate 68.9% of the time. A schema check cannot see that gap, which is the entire point of scoring both.
- Let ambiguous inputs have ambiguous labels
- Forcing one answer onto a ticket that legitimately belongs to two teams measures the labeller rather than the model, so the label set allows more than one acceptable answer.
What the measurements showed
- Enforcement tier barely moved structural success; model capacity decided it. On no model did enforcement beat the prompt-only baseline.
- Schema complexity predicted failure: one model scored 86.7% flat, 16.7% nested, and 0.0% on the hard schema across 60 attempts.
- Valid is not correct: 100% structurally valid, 68.9% semantically accurate.
What broke
- An earlier version of Customer.name was a required non-nullable string. Eight of the ten test tickets name nobody, so models had no legal way to say "not stated", and the benchmark recorded 80% of extractions as inventing a customer name. They were not inventing: 292 of the flagged values were the literal string the field description had asked for. With the field nullable and the description asking for null, grounding is 100% on every model except the 7B one. Give a model a way to decline and it takes it.
- The LLM critic scoring semantic accuracy originally agreed with hand labels only 50% of the time, which is chance. It could not see the allowed enum values and treated inferred judgements like priority as facts to be found in the text. That version would have published a 91.7% failure rate for a system whose real rate was 41.7%. An unmeasured judge is just a second unvalidated model.
Running it
pip install -r requirements.txtFull setup, configuration and API reference are in the repository README.
Stack
- Python
- Pydantic
- SQLite
- Groq
- OpenRouter