Reliability and Evaluation Research · Feb – Aug 2026
ragwell: RAG with Verified Citation Grounding
126× token cost difference between retrieval and long context
The problem
Every AI tool shows sources now. Almost none verify the sources actually say what the answer claims.
Approach
A claim type that cannot exist without a chunk id and a verbatim quote, which makes citations checkable by string containment: no second model, no cost, and it runs on every answer rather than on a sample. Five search strategies from keyword matching to hypothetical document expansion. A three-rung verification ladder: quote containment, lexical overlap, entailment. InsufficientEvidence is a first-class answer, because a system with no way to say that will say something else instead. 381 tests, benchmarked on four real regulatory annual reports.
How it works
- A claim cannot exist without a chunk id and a verbatim quote, which makes every citation checkable by string containment.
- Verification runs in three rungs, each reported separately so the cheap ones can be counted alone: quote containment, lexical overlap, then entailment.
- Quote containment is whitespace-normalised, because models reflow text when they quote it.
- The entailment judge is a model judging a model, so its agreement with hand labels is measured before any number it produces is quoted: 8 of 8, zero false supports, stable across three runs.
- InsufficientEvidence is a first-class answer, because a system with no way to decline will say something else instead.
Key decisions
- Verification that costs nothing runs on everything
- String containment is free and deterministic, so it runs on every answer rather than on a sample. Only the entailment rung costs a model call, and it sits last precisely because it is the expensive one.
- Publish the calibration curve, not a threshold
- The confidence score is not a probability: ECE is 15.1%. No threshold paid for itself on this corpus, so the curve is published rather than a number that would imply a precision the score does not have.
- Quote the limits at their real size
- Eight labelled cases is small and is described that way. Ten questions per cell, one regulator's annual reports, dense and numeric and English. Strong enough for the cost and trap findings, not for accuracy differences.
Interface

What the measurements showed
- Retrieval held flat at ~2,600 tokens per question while long context ran to 319,771. Arithmetic, not sampling.
- Long context did not solve superseded documents, stating outdated figures on 67% of misleading questions against retrieval's 33%.
- Chunk size bought 7.9× citation precision for free, with accuracy flat.
- The confidence score is not a probability: ECE 15.1%. No threshold paid for itself, so the curve is published instead of a number.
What broke
- A substring match on money, where N5,000,000 contains 500,000, scored every correct answer to the central question as wrong. It would have inverted the headline.
- A partial benchmark run suggested the opposite conclusion on superseded documents and was nearly published.
Running it
python -m uvicorn app.main:app --port 8000Full setup, configuration and API reference are in the repository README.
Stack
- Python
- FastAPI
- ChromaDB
- Jina AI
- Gemini
- Groq
- Next.js