From 72% to 95%: fixing an assistant that only quotes its sources

A company knowledge assistant that answers only in sentences quoted from its documents went from 72% to 95% correct on a 120-question test, without loosening a single scoring rule.

Update, 25 September 2026: a later change took the assistant to 118 of 120. The follow-up is From 116 to 118 of 120: why I threw away the fix that scored perfectly.

Useful Brain is a question-answering assistant I built for company documents: handbooks, policies, plans. An employee asks a question and gets an answer made only of sentences that actually appear in those documents, each with a citation, or a plain “not enough evidence” when the documents do not say. That matters because a confident wrong answer about a leave policy or an expense rule is worse than no answer. The same rule is what made it hard to test: every sentence in an answer must be an exact span of a retrieved document, cited with a label, or the host software rejects the whole answer.

This is the record of two repair passes that moved the assistant’s live test score from 77 of 107 (72%, with 13 questions skipped) to 114 of 120 (95%). The test, an eval in the jargon, is a fixed set of 120 questions with known right answers and known right documents, scored automatically. Multi-hop questions, the ones whose answer needs two documents, went from 0 of 10 to 9 of 10. Every gain came from the answer layer. The retrieval step that finds candidate passages, and every scoring rule, stayed locked the whole time.

Live pass rateNorthwind battery, 120 locked questions, three runs

95.0%final pass rate, up from 72.0%, with every scoring rule locked

Live pass rate on the Northwind battery. The baseline number looks better than it was: it excludes the 13 permission questions the live path could not run.

The setup

The corpus is 65 synthetic company documents with document-level access rules (public, department, role, private owner). Each question is asked as a particular person, its principal, who is only allowed to see some documents. The battery is 120 locked questions:

  • 70 factual questions
  • 17 traps, where a plausible wrong document sits next to the right one
  • 13 permission cases, where the answer lives in a document the asking principal cannot read
  • 10 unanswerable questions
  • 10 multi-hop questions whose answer requires citing two documents

The scoring rules are harsh on purpose and were never loosened:

  • A multi-hop answer must cite every gold document. One of two is a fail.
  • Permission and unanswerable questions must return insufficient_evidence. A fluent answer from a related-but-wrong document is a fail.
  • Retrieving any forbidden document is a fail regardless of the answer.
  • Every question runs under its own principal, and the harness fails closed if the backend does not confirm the identity or the answering model.

The answering stack is hybrid retrieval (candidate passages found by keyword and by meaning, then re-ordered by a second model) feeding a language model on Cloudflare Workers AI that can call a search tool, with a host-side grounding validator that only accepts answers whose every sentence is a verbatim span of retrieved evidence.

The answer pipelineWhere each fix landed

  1. Question under the asking principalPass 1Loopback-only assumed principal so each question runs as its own user
  2. Hybrid retrieval and ACL filterLockedKeyword, vector, rerank. Unchanged in both passes
  3. Model draft with search toolFix B, C, DDocument identity before text, twin-document rule, temperature 0 and pinned seed
  4. Citation completionFix AThinking disabled, JSON recovered from narrated output
  5. Coverage pass, multi-part questions onlyFix AAsks for the exact evidence sentence answering each open part
  6. Refusal guard, then verbatim salvageFix D, reviewRebuilds an invalid draft from its exact spans. Order fixed after review
  7. Grounding validatorLockedEvery sentence verbatim and cited, or insufficient_evidence
The answer pipeline. Retrieval and the validator never changed; every fix sits between them.

Pass 1: score honestly first

The original 72% hid two measurement problems. Permission questions were skipped on the live path because the local operator identity could not represent the asking principal. And several factual failures were really identity failures: role-scoped gold documents never entered retrieval for the loopback operator.

Pass 1 fixed the measurement before the model. A loopback-only assumed-principal field scoped retrieval to each question’s principal (never storage or tool policy, and rejected outside loopback). Deterministic citation completion attached labels the model had earned but not written. Refusal handling stopped routing genuine abstentions through citation repair.

Score: 95 of 120, with all 120 finally scored. But multi-hop collapsed to 0 of 10, and four failure families remained.

Pass 1, one cell per question95 of 120 passed

0/10multi-hop questions passed after Pass 1

Pass 1. Each cell is one question; highlighted cells failed. Both multi-hop rows are entirely highlighted.

Pass 2: four failure families, four root causes

A. Multi-hop 0/10: second document retrieved, never cited

Retrieval delivered both gold documents on nine of the ten multi-hop questions (mean gold-document recall 0.95). Then the model wrote about one of them.

The real root cause was invisible until I logged the raw model responses. The citation-repair and coverage calls capped completion at 512 tokens, and GLM 5.3 Flash is a reasoning model. It spent the entire budget thinking, hit finish_reason: length, and returned empty content. Every “repair” was silently a no-op, falling back to a crude lexical extractor.

The 512-token budget on a repair callBefore and after disabling thinking

512/512tokens spent on reasoning, none on the answer

What a 512-token cap does to a reasoning model asked to select quotes. The call returned finish_reason: length and an empty body on every repair.

Fixes: disable thinking on extraction calls (they select quotes; they do not need chain-of-thought). Recover the JSON object from narrated responses with balanced-brace parsing that prefers the last emitted object over mid-reasoning examples. Add a coverage pass that asks, for multi-part questions only, for the exact evidence sentence answering each still-open part. Coverage additions must re-validate against the evidence ledger before they are kept.

B. Twin-document misattribution

Six factual questions failed because the model quoted the lookalike sentence from a neighbouring document (the handbook instead of the dedicated policy). Fixes: each search hit now leads with its document identity before the text, and the prompt prefers the dedicated policy document for the asked topic, or cites both.

C. Grounded answers from allowed neighbours on permission questions

The forbidden document was correctly never retrieved, but the model answered anyway from a related, allowed document. The fix was discipline, not machinery: the prompt now states that a sentence about a different programme, plan or policy than the one asked is not an answer.

D. Run-to-run churn

Questions flipped between runs. Fixes: temperature 0 and a pinned seed on every call, plus a verbatim-salvage pass that deterministically rebuilds an invalid draft from its exact evidence spans. The model kept wrapping correct quotes in **Label:** "quote" - from file.md decoration that failed strict validation; salvage strips the decoration and keeps the span.

Pass 1 to Pass 2, by categoryHollow is Pass 1, filled is Pass 2

0 → 9multi-hop questions of 10

Pass 1 (hollow) to Pass 2 (filled), as a share of each category. Multi-hop is the whole story of family A.

The adversarial review earned its keep

Before committing, three parallel reviewers (security, grounding logic, eval honesty) attacked the diff. The security reviewer found a real high-severity bug: the new salvage pass ran ahead of the refusal-honouring guard, so a short refusal that happened to quote an evidence sentence could be converted into a confident grounded answer to an unanswerable question. An interim run had scored 115 of 120 with that bug in place.

I fixed the order, hardened salvage three more ways (body-text-only grounding, longest-span preference, bounded colon prefixes), tightened the multi-part trigger so a narrative “and” in trap questions cannot fire it, and re-ran everything. The recorded 114 of 120 is the honest post-fix number, one point lower than the flattering one.

The eval harness also got integrity upgrades from the review: checkpoints pin the exact answer-pipeline build and a digest of the question set, resumes fail closed on any mismatch, latency summaries carry a partial flag, and the findings report a cited-but-not-expected counter so citation broadening cannot hide. For the final run that counter reads 9 of 120 grounded answers citing a document outside the gold set, and gold documents retrieved but uncited dropped from 31 rows in Pass 1 to 16.

Pass 2, one cell per question114 of 120 passed

6failures remain, four of them fail for every model tested

Pass 2, the recorded final run. Six failures remain: q028, q073, q088, q100, q105, q110.

What stayed constant

The deterministic in-process retrieval suite (fake embeddings, model-independent, run as a regression guard beside every live pass) held at recall@3 0.912, MRR 0.825 and nDCG 0.837 with zero ACL leaks throughout. It measures the retrieval code, not the live Cloudflare stack. The live-stack retrieval signal is a retrieved recall of 0.974 in the final run, with zero forbidden-document retrievals across every pass.

Honest caveats

  • Three to five questions still churn between runs. Temperature 0 does not pin a reasoning model’s thinking on this stack, so treat single-question flips as noise and category totals as signal.
  • Complete-answer latency (p50 about 15.6s, p95 about 56s) exceeds the 15s p95 production target. Most of that is the model’s reasoning phase, and the final run executed beside four other model runs on one local worker, so these numbers include contention. An earlier solo segment measured p50 about 19.6s. Recorded as an open item, not accepted.
  • Six failures remain. Four of them (q073, q100, q105, q110) fail for every model I later tested, which points at corpus difficulty and abstention discipline rather than model choice. The other two are honest residue of the repaired families: q028 is a twin-document misattribution (family B improved, not closed) and q088 dropped a multi-hop second citation (family A at 9 of 10, not 10).

Evidence

The frozen findings for both passes live in the repo: findings.pass1.json and findings.glm-5.3-flash.json. The 30 August baseline (77 of 107) is recorded in the execution tracker and has no frozen snapshot. The reproduce commands are in the evals index.

The follow-up post runs the same battery against every plausible Cloudflare-hosted chat model: Picking a model for a grounded RAG agent.