Skip to content

Reliability and Evaluation Research · Aug 2026

so-agent: Structured Output Reliability Benchmark

87.6% single-call success collapses to 26.7% over ten chained calls

Source

The problem

Everyone assumes structured output enforcement works. Nobody measures whether providers actually honour it.

Approach

A capability probe distinguishing four outcomes per enforcement tier. The critical one is IGNORED, where a provider accepts a directive and silently does not apply it. A three-tier enforcement ladder selected from the probe. A typed failure taxonomy separating free repairs from unretryable refusals. Every rate published with its confidence interval and n, across 833 tests on 6 models, backed by a 340-test suite.

How it works

Key decisions

Probe the provider, do not trust the documentation
Every cell in the capability matrix is a real API call. Three of six models on one provider reject a native JSON schema outright, which is why the ladder exists rather than being a formality: on half the line-up there is nothing to fall back from.
Measure structure and meaning separately
gpt-oss-20b produces structurally valid output 100% of the time at every tier and is semantically accurate 68.9% of the time. A schema check cannot see that gap, which is the entire point of scoring both.
Let ambiguous inputs have ambiguous labels
Forcing one answer onto a ticket that legitimately belongs to two teams measures the labeller rather than the model, so the label set allows more than one acceptable answer.

What the measurements showed

What broke

Running it

pip install -r requirements.txt

Full setup, configuration and API reference are in the repository README.

Stack

Written up in