# From 72% to 95%: fixing an assistant that only quotes its sources

> A company knowledge assistant that answers only in sentences quoted from its documents went from 72% to 95% correct on a 120-question test, without loosening a single scoring rule.

- Author: Wasim Jalali
- Published: 2026-09-01
- URL: https://wasimjalali.com/writings/from-72-to-95-repairing-a-grounded-rag-agent/

*Update, 25 September 2026: a later change took the assistant to 118 of 120. The follow-up is [From 116 to 118 of 120: why I threw away the fix that scored perfectly](/writings/from-116-to-118-of-120-the-fix-i-threw-away/).*

[Useful Brain](https://github.com/wasimjalali/useful-brain) is a question-answering assistant I built for company documents: handbooks, policies, plans. An employee asks a question and gets an answer made only of sentences that actually appear in those documents, each with a citation, or a plain "not enough evidence" when the documents do not say. That matters because a confident wrong answer about a leave policy or an expense rule is worse than no answer. The same rule is what made it hard to test: every sentence in an answer must be an exact span of a retrieved document, cited with a label, or the host software rejects the whole answer.

This is the record of two repair passes that moved the assistant's live test score from 77 of 107 (72%, with 13 questions skipped) to 114 of 120 (95%). The test, an eval in the jargon, is a fixed set of 120 questions with known right answers and known right documents, scored automatically. Multi-hop questions, the ones whose answer needs two documents, went from 0 of 10 to 9 of 10. Every gain came from the answer layer. The retrieval step that finds candidate passages, and every scoring rule, stayed locked the whole time.

<figure class="fig is-field">
<div class="fig-top">
<p class="fig-head">Live pass rate<small>Northwind battery, 120 locked questions, three runs</small></p>
<p class="fig-callout"><b>95.0%</b>final pass rate, up from 72.0%, with every scoring rule locked</p>
</div>
<div class="bars" role="img" aria-label="Pass rate across three runs: 72%, 79.2%, 95.0%">
<div class="bar-row"><span class="bar-label">Baseline, 30 Aug<small>13 permission questions skipped</small></span><span class="bar-track"><span class="bar-fill" style="width:72%"></span></span><span class="bar-value">77/107<b>72.0%</b></span></div>
<div class="bar-row"><span class="bar-label">Pass 1, 31 Aug<small>all 120 scored</small></span><span class="bar-track"><span class="bar-fill" style="width:79.2%"></span></span><span class="bar-value">95/120<b>79.2%</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Pass 2, final<small>all 120 scored</small></span><span class="bar-track"><span class="bar-fill" style="width:95%"></span></span><span class="bar-value">114/120<b>95.0%</b></span></div>
</div>
<figcaption>Live pass rate on the Northwind battery. The baseline number looks better than it was: it excludes the 13 permission questions the live path could not run.</figcaption>
</figure>

## The setup

The corpus is 65 synthetic company documents with document-level access rules (public, department, role, private owner). Each question is asked as a particular person, its principal, who is only allowed to see some documents. The battery is 120 locked questions:

- 70 factual questions
- 17 traps, where a plausible wrong document sits next to the right one
- 13 permission cases, where the answer lives in a document the asking principal cannot read
- 10 unanswerable questions
- 10 multi-hop questions whose answer requires citing two documents

The scoring rules are harsh on purpose and were never loosened:

- A multi-hop answer must cite every gold document. One of two is a fail.
- Permission and unanswerable questions must return `insufficient_evidence`. A fluent answer from a related-but-wrong document is a fail.
- Retrieving any forbidden document is a fail regardless of the answer.
- Every question runs under its own principal, and the harness fails closed if the backend does not confirm the identity or the answering model.

The answering stack is hybrid retrieval (candidate passages found by keyword and by meaning, then re-ordered by a second model) feeding a language model on Cloudflare Workers AI that can call a search tool, with a host-side grounding validator that only accepts answers whose every sentence is a verbatim span of retrieved evidence.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">The answer pipeline<small>Where each fix landed</small></p>
</div>
<ol class="flow" aria-label="Answer pipeline and where each fix landed">
<li><span class="flow-name">Question under the asking principal</span><span class="flow-tag">Pass 1</span><span class="flow-note">Loopback-only assumed principal so each question runs as its own user</span></li>
<li class="is-locked"><span class="flow-name">Hybrid retrieval and ACL filter</span><span class="flow-tag">Locked</span><span class="flow-note">Keyword, vector, rerank. Unchanged in both passes</span></li>
<li><span class="flow-name">Model draft with search tool</span><span class="flow-tag">Fix B, C, D</span><span class="flow-note">Document identity before text, twin-document rule, temperature 0 and pinned seed</span></li>
<li><span class="flow-name">Citation completion</span><span class="flow-tag">Fix A</span><span class="flow-note">Thinking disabled, JSON recovered from narrated output</span></li>
<li><span class="flow-name">Coverage pass, multi-part questions only</span><span class="flow-tag">Fix A</span><span class="flow-note">Asks for the exact evidence sentence answering each open part</span></li>
<li><span class="flow-name">Refusal guard, then verbatim salvage</span><span class="flow-tag">Fix D, review</span><span class="flow-note">Rebuilds an invalid draft from its exact spans. Order fixed after review</span></li>
<li class="is-locked"><span class="flow-name">Grounding validator</span><span class="flow-tag">Locked</span><span class="flow-note">Every sentence verbatim and cited, or <code>insufficient_evidence</code></span></li>
</ol>
<figcaption>The answer pipeline. Retrieval and the validator never changed; every fix sits between them.</figcaption>
</figure>

## Pass 1: score honestly first

The original 72% hid two measurement problems. Permission questions were skipped on the live path because the local operator identity could not represent the asking principal. And several factual failures were really identity failures: role-scoped gold documents never entered retrieval for the loopback operator.

Pass 1 fixed the measurement before the model. A loopback-only assumed-principal field scoped retrieval to each question's principal (never storage or tool policy, and rejected outside loopback). Deterministic citation completion attached labels the model had earned but not written. Refusal handling stopped routing genuine abstentions through citation repair.

Score: 95 of 120, with all 120 finally scored. But multi-hop collapsed to 0 of 10, and four failure families remained.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Pass 1, one cell per question<small>95 of 120 passed</small></p>
<p class="fig-callout"><b>0/10</b>multi-hop questions passed after Pass 1</p>
</div>
<div class="grid120" role="img" aria-label="Pass 1: 95 of 120 questions passed, 25 failed">
<div class="cg"><span class="cg-label">Factual <b>59/70</b></span><span class="cells"><i class="fail" title="q001"></i><i></i><i></i><i class="fail" title="q004"></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q024"></i><i></i><i class="fail" title="q026"></i><i></i><i class="fail" title="q028"></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q038"></i><i></i><i class="fail" title="q040"></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q093"></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q100"></i><i></i><i></i><i></i><i></i><i class="fail" title="q105"></i><i></i><i></i><i></i><i></i><i class="fail" title="q110"></i></span></div>
<div class="cg"><span class="cg-label">Trap <b>17/17</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Permission <b>10/13</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q072"></i><i class="fail" title="q073"></i><i class="fail" title="q074"></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Unanswerable <b>9/10</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q085"></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop <b>0/5</b></span><span class="cells"><i class="fail" title="q086"></i><i class="fail" title="q087"></i><i class="fail" title="q088"></i><i class="fail" title="q089"></i><i class="fail" title="q090"></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop, expanded <b>0/5</b></span><span class="cells"><i class="fail" title="q116"></i><i class="fail" title="q117"></i><i class="fail" title="q118"></i><i class="fail" title="q119"></i><i class="fail" title="q120"></i></span></div>
</div>
<figcaption>Pass 1. Each cell is one question; highlighted cells failed. Both multi-hop rows are entirely highlighted.</figcaption>
</figure>

## Pass 2: four failure families, four root causes

### A. Multi-hop 0/10: second document retrieved, never cited

Retrieval delivered both gold documents on nine of the ten multi-hop questions (mean gold-document recall 0.95). Then the model wrote about one of them.

The real root cause was invisible until I logged the raw model responses. The citation-repair and coverage calls capped completion at 512 tokens, and GLM 5.3 Flash is a reasoning model. It spent the entire budget thinking, hit `finish_reason: length`, and returned empty content. Every "repair" was silently a no-op, falling back to a crude lexical extractor.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">The 512-token budget on a repair call<small>Before and after disabling thinking</small></p>
<p class="fig-callout"><b>512/512</b>tokens spent on reasoning, none on the answer</p>
</div>
<div class="budget" role="img" aria-label="Before: the 512-token budget was fully consumed by reasoning and no content was returned. After: reasoning disabled, the JSON answer fits.">
<div class="budget-row"><span class="budget-label">Before</span><span class="budget-track"><span class="budget-think" style="width:100%">reasoning</span></span><span class="budget-value">512/512 used<b>content: empty</b></span></div>
<div class="budget-row"><span class="budget-label">After</span><span class="budget-track"><span class="budget-out" style="width:22%">JSON</span></span><span class="budget-value">thinking off<b>quotes returned</b></span></div>
</div>
<figcaption>What a 512-token cap does to a reasoning model asked to select quotes. The call returned <code>finish_reason: length</code> and an empty body on every repair.</figcaption>
</figure>

Fixes: disable thinking on extraction calls (they select quotes; they do not need chain-of-thought). Recover the JSON object from narrated responses with balanced-brace parsing that prefers the last emitted object over mid-reasoning examples. Add a coverage pass that asks, for multi-part questions only, for the exact evidence sentence answering each still-open part. Coverage additions must re-validate against the evidence ledger before they are kept.

### B. Twin-document misattribution

Six factual questions failed because the model quoted the lookalike sentence from a neighbouring document (the handbook instead of the dedicated policy). Fixes: each search hit now leads with its document identity before the text, and the prompt prefers the dedicated policy document for the asked topic, or cites both.

### C. Grounded answers from allowed neighbours on permission questions

The forbidden document was correctly never retrieved, but the model answered anyway from a related, allowed document. The fix was discipline, not machinery: the prompt now states that a sentence about a different programme, plan or policy than the one asked is not an answer.

### D. Run-to-run churn

Questions flipped between runs. Fixes: temperature 0 and a pinned seed on every call, plus a verbatim-salvage pass that deterministically rebuilds an invalid draft from its exact evidence spans. The model kept wrapping correct quotes in `**Label:** "quote" - from file.md` decoration that failed strict validation; salvage strips the decoration and keeps the span.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Pass 1 to Pass 2, by category<small>Hollow is Pass 1, filled is Pass 2</small></p>
<p class="fig-callout"><b>0 → 9</b>multi-hop questions of 10</p>
</div>
<div class="dumbbells" role="img" aria-label="Per-category pass counts, Pass 1 to Pass 2: factual 59 to 66 of 70, trap 17 to 17, permission 10 to 12 of 13, unanswerable 9 to 10, multi-hop 0 to 9 of 10">
<div class="db-row"><span class="db-label">Factual</span><span class="db-track"><span class="db-seg" style="left:84.3%;width:10%"></span><span class="db-a" style="left:84.3%"></span><span class="db-b" style="left:94.3%"></span></span><span class="db-value">59 → 66 <small>of 70</small></span></div>
<div class="db-row"><span class="db-label">Trap</span><span class="db-track"><span class="db-b" style="left:100%"></span></span><span class="db-value">17 → 17 <small>of 17</small></span></div>
<div class="db-row"><span class="db-label">Permission</span><span class="db-track"><span class="db-seg" style="left:76.9%;width:15.4%"></span><span class="db-a" style="left:76.9%"></span><span class="db-b" style="left:92.3%"></span></span><span class="db-value">10 → 12 <small>of 13</small></span></div>
<div class="db-row"><span class="db-label">Unanswerable</span><span class="db-track"><span class="db-seg" style="left:90%;width:10%"></span><span class="db-a" style="left:90%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">9 → 10 <small>of 10</small></span></div>
<div class="db-row"><span class="db-label">Multi-hop</span><span class="db-track"><span class="db-seg" style="left:0%;width:90%"></span><span class="db-a" style="left:0%"></span><span class="db-b" style="left:90%"></span></span><span class="db-value">0 → 9 <small>of 10</small></span></div>
</div>
<figcaption>Pass 1 (hollow) to Pass 2 (filled), as a share of each category. Multi-hop is the whole story of family A.</figcaption>
</figure>

## The adversarial review earned its keep

Before committing, three parallel reviewers (security, grounding logic, eval honesty) attacked the diff. The security reviewer found a real high-severity bug: the new salvage pass ran ahead of the refusal-honouring guard, so a short refusal that happened to quote an evidence sentence could be converted into a confident grounded answer to an unanswerable question. An interim run had scored 115 of 120 with that bug in place.

I fixed the order, hardened salvage three more ways (body-text-only grounding, longest-span preference, bounded colon prefixes), tightened the multi-part trigger so a narrative "and" in trap questions cannot fire it, and re-ran everything. The recorded 114 of 120 is the honest post-fix number, one point lower than the flattering one.

The eval harness also got integrity upgrades from the review: checkpoints pin the exact answer-pipeline build and a digest of the question set, resumes fail closed on any mismatch, latency summaries carry a partial flag, and the findings report a cited-but-not-expected counter so citation broadening cannot hide. For the final run that counter reads 9 of 120 grounded answers citing a document outside the gold set, and gold documents retrieved but uncited dropped from 31 rows in Pass 1 to 16.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Pass 2, one cell per question<small>114 of 120 passed</small></p>
<p class="fig-callout"><b>6</b>failures remain, four of them fail for every model tested</p>
</div>
<div class="grid120" role="img" aria-label="Pass 2: 114 of 120 questions passed, 6 failed">
<div class="cg"><span class="cg-label">Factual <b>66/70</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q028"></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q100"></i><i></i><i></i><i></i><i></i><i class="fail" title="q105"></i><i></i><i></i><i></i><i></i><i class="fail" title="q110"></i></span></div>
<div class="cg"><span class="cg-label">Trap <b>17/17</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Permission <b>12/13</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q073"></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Unanswerable <b>10/10</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop <b>4/5</b></span><span class="cells"><i></i><i></i><i class="fail" title="q088"></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop, expanded <b>5/5</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i></span></div>
</div>
<figcaption>Pass 2, the recorded final run. Six failures remain: q028, q073, q088, q100, q105, q110.</figcaption>
</figure>

## What stayed constant

The deterministic in-process retrieval suite (fake embeddings, model-independent, run as a regression guard beside every live pass) held at recall@3 0.912, MRR 0.825 and nDCG 0.837 with zero ACL leaks throughout. It measures the retrieval code, not the live Cloudflare stack. The live-stack retrieval signal is a retrieved recall of 0.974 in the final run, with zero forbidden-document retrievals across every pass.

## Honest caveats

- Three to five questions still churn between runs. Temperature 0 does not pin a reasoning model's thinking on this stack, so treat single-question flips as noise and category totals as signal.
- Complete-answer latency (p50 about 15.6s, p95 about 56s) exceeds the 15s p95 production target. Most of that is the model's reasoning phase, and the final run executed beside four other model runs on one local worker, so these numbers include contention. An earlier solo segment measured p50 about 19.6s. Recorded as an open item, not accepted.
- Six failures remain. Four of them (q073, q100, q105, q110) fail for every model I later tested, which points at corpus difficulty and abstention discipline rather than model choice. The other two are honest residue of the repaired families: q028 is a twin-document misattribution (family B improved, not closed) and q088 dropped a multi-hop second citation (family A at 9 of 10, not 10).

## Evidence

The frozen findings for both passes live in the repo: [`findings.pass1.json`](https://github.com/wasimjalali/useful-brain/blob/main/evals/results/2026-08-31/findings.pass1.json) and [`findings.glm-5.3-flash.json`](https://github.com/wasimjalali/useful-brain/blob/main/evals/results/2026-08-31/findings.glm-5.3-flash.json). The 30 August baseline (77 of 107) is recorded in the execution tracker and has no frozen snapshot. The reproduce commands are in the [evals index](https://github.com/wasimjalali/useful-brain/tree/main/evals).

The follow-up post runs the same battery against every plausible Cloudflare-hosted chat model: [Picking a model for a grounded RAG agent](/writings/picking-a-model-for-a-grounded-rag-agent/).
