# From 116 to 118 of 120: why I threw away the fix that scored perfectly

> A company knowledge assistant that only quotes its documents reached 118 of 120 (98.3%) on its live test. The first fix repaired all four misses on replay and was discarded anyway, because it only worked for this set of documents.

- Author: Wasim Jalali
- Published: 2026-09-25
- URL: https://wasimjalali.com/writings/from-116-to-118-of-120-the-fix-i-threw-away/

[Useful Brain](https://github.com/wasimjalali/useful-brain) is a question-answering assistant I built for company documents: handbooks, policies, plans. An employee asks a question and gets an answer made only of sentences that actually appear in those documents, each with a citation, or a plain "not enough evidence" when the documents do not say. After the [repair work in August](/writings/from-72-to-95-repairing-a-grounded-rag-agent/) it scored 114 of 120 on its live test. By September the misses that kept coming back had one shape: the assistant found the right document and then quoted a lookalike sentence from the document next to it.

This post is about the day I fixed that twice. A fresh run on the morning of 6 September scored 116 of 120. The first fix I wrote repaired all four misses on replay, and I deleted it, because it worked by matching this corpus's wording and would not have survived a new set of documents. The fix I kept is smaller and more general, and took the run to 118 of 120 (98.3%). The honest framing: runs of this system land anywhere between 116 and 120, so treat the two-point move as the category totals moving, not a pinned number.

<figure class="fig is-field">
<div class="fig-top">
<p class="fig-head">One morning, two fixes<small>Northwind battery, 120 locked questions, GLM 5.3 Flash</small></p>
<p class="fig-callout"><b>118/120</b>98.3% with the general fix, the strongest run on this battery</p>
</div>
<div class="fork" role="img" aria-label="A fresh run scored 116 of 120 with four misses of the same shape. The first fix, host-side sentence grafting, repaired 4 of 4 on replay and was discarded because it fitted this corpus's wording. The second fix, pointer detection with the model doing the wording, was kept and scored 118 of 120.">
<div class="fork-root"><span class="fork-kicker">Fresh run, 6 September</span><b>116/120</b><span class="fork-note">Four misses, one shape: the right documents were retrieved, the answer quoted a neighbouring twin</span></div>
<div class="fork-branches">
<div class="fork-branch is-dropped"><span class="fork-kicker">Fix 1, discarded</span><b>4/4</b><span class="fork-note">The host grafts the missing sentence by word overlap. Repaired every miss on replay. Deleted: it matched this corpus's phrasing, not the problem</span></div>
<div class="fork-branch is-focal"><span class="fork-kicker">Fix 2, kept</span><b>118/120</b><span class="fork-note">The host detects which document the evidence points at; the model writes the answer. No question ids, no corpus words in the code</span></div>
</div>
</div>
<figcaption>The fix that scored 4 of 4 on the misses was the one thrown away. Strength over score.</figcaption>
</figure>

## What was tested

The test, an eval in the jargon, is a fixed set of 120 questions with known right answers and known right documents, scored automatically: 70 factual questions, 17 traps where a plausible wrong document sits next to the right one, 13 permission cases where the answer lives in a document the asking person cannot read, 10 unanswerable questions and 10 multi-hop questions whose answer needs two documents. The corpus is 65 synthetic company documents with access rules. Every scoring rule is the same as in August and none was loosened: every sentence must be a verbatim span of a retrieved document, a multi-hop answer must cite both documents, and a permission or unanswerable question must come back as `insufficient_evidence`.

The model was GLM 5.3 Flash on Cloudflare Workers AI, the one chosen in [the model bake-off](/writings/picking-a-model-for-a-grounded-rag-agent/). Retrieval, the step that finds candidate passages, was locked for the whole of this work. Only the answer layer, the code between retrieval and the validator that checks every sentence, was allowed to change.

## Four misses, one shape

The fresh run's four misses were one problem wearing four outfits. In each, retrieval delivered the gold documents, the ones the answer should cite. The draft then cited a neighbouring document instead: the handbook restating a rule the dedicated policy owns, or a page that mentions a programme it does not define.

Company documents do this constantly. A retention section says "request earlier deletion through the Customer Data Access Requests process". A recruiting page says "the hiring bonus is a different program". The sentence in front of the model names the document that owns the answer, and the model, having found a sentence that sounds right, stops there. I started calling these pointer cases: the evidence itself points somewhere else.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">A pointer case<small>What the model saw on each of the four misses</small></p>
</div>
<div class="pointer" role="img" aria-label="The cited document contains a sentence that names another document by title. That other document was retrieved but not cited, and it is the one that owns the answer.">
<div class="pointer-card"><span class="pointer-kicker">Cited by the draft</span><span class="pointer-title">A neighbouring document</span><span class="pointer-quote">"request earlier deletion through the <mark>Customer Data Access Requests</mark> process"</span></div>
<div class="pointer-link"><span>names it by title</span></div>
<div class="pointer-card is-focal"><span class="pointer-kicker">Retrieved, not cited</span><span class="pointer-title">Customer Data Access Requests</span><span class="pointer-quote">The document that owns the rule. It was in the evidence the whole time.</span></div>
</div>
<figcaption>Both documents were in front of the model. The cited one only points at the answer; the uncited one holds it.</figcaption>
</figure>

## The fix that worked and got deleted

My first fix was host-side sentence grafting. When a draft cited a neighbour and left an owner document uncited, the host itself picked a sentence from the owner, choosing by how many words it shared with the question, and grafted it into the answer before validation. I replayed the four failing questions and it repaired all four. On the scoreboard it was finished.

It was also wrong, and I knew it the moment I read the selection rule back. Word overlap picks the right sentence in this corpus because these 65 documents were generated with consistent phrasing: the question and the owning sentence tend to share vocabulary. Point the same code at a real company's documents, where the question says "days off" and the policy says "annual leave entitlement", and the graft picks a wrong sentence with the same confidence. Worse, a wrong graft would arrive already stitched into an answer that reads fluently. The whole reason this assistant exists is that a confident wrong answer about a leave policy is worse than no answer.

So the grafting went in the bin, with its 4 of 4. The report records it under "deliberately not done". The lesson I keep relearning on this project is that a fix that fits the test is a fix that fits the test, and the number it produces is not evidence of anything except that.

## The fix I kept

The kept design splits the job in two. The host does the part that can be done deterministically and without any knowledge of the corpus: noticing that the evidence points at an uncited document. The model does the part that needs judgement: deciding whether that document actually answers the question, and in which sentence. Three changes, all in the answer layer, shipped as prompt version `grounded-answer.v9`.

**Pointer detection.** A small module reads the question, the draft and the sentences the draft cited, and looks for uncited evidence documents named by their title words, plus documents whose text explicitly contrasts itself with a cited one ("is a different program", "separate from", "not interchangeable", "does not replace"). It contains no question ids and no corpus vocabulary. It selects no sentences. Its only output is a hint: these documents were named, take a second look.

**A wider coverage pass.** The assistant already had a coverage pass, a second model call that asks, for multi-part questions, for the exact evidence sentence answering each still-open part. It now also runs for single-part drafts that carry a pointer hint, with a stronger instruction about telling similar programmes apart. That is one extra model call, only when the evidence itself points elsewhere. As before, the host keeps an addition only if it re-validates against the evidence ledger, so the model cannot smuggle in a sentence that is not there.

**Two prompt rules, and a bug.** The chat prompt now says that when the evidence calls two programmes different, each must be answered from its own document. And while wiring the hints I found that single-word titles (`privacy-policy.md` reduces to one distinctive word once "policy" is dropped as a stop word) could never trigger a coverage hint under the old rule, which demanded two matching title words. The threshold is now the smaller of two and the number of words the title has.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">The answer pipeline<small>Where the September change landed</small></p>
</div>
<ol class="flow" aria-label="Answer pipeline. Retrieval and the grounding validator are locked. The change adds pointer detection before a widened coverage pass, plus a prompt rule on the model draft.">
<li class="is-locked"><span class="flow-name">Hybrid retrieval and access filter</span><span class="flow-tag">Locked</span><span class="flow-note">Keyword, vector, rerank. Not touched in this work</span></li>
<li><span class="flow-name">Model draft with search tool</span><span class="flow-tag">Prompt v9</span><span class="flow-note">Programmes the evidence calls different are answered each from its own document</span></li>
<li><span class="flow-name">Pointer detection</span><span class="flow-tag">New</span><span class="flow-note">Uncited documents named by title, or contrasting themselves with a cited one. Detection only; no sentences chosen</span></li>
<li><span class="flow-name">Coverage pass</span><span class="flow-tag">Widened</span><span class="flow-note">Was multi-part questions only. Now also single-part drafts with a pointer hint. Additions must re-validate</span></li>
<li class="is-locked"><span class="flow-name">Refusal guard, then verbatim salvage</span><span class="flow-tag">Unchanged</span><span class="flow-note">As left after the August review</span></li>
<li class="is-locked"><span class="flow-name">Grounding validator</span><span class="flow-tag">Locked</span><span class="flow-note">Every sentence verbatim and cited, or <code>insufficient_evidence</code></span></li>
</ol>
<figcaption>The host gained one deterministic step that points; the model kept the step that decides.</figcaption>
</figure>

## 116 to 118

The full battery ran again with the kept fix. Factual, trap and permission questions did not move, because they were already at or within one of their ceilings. Two categories moved: unanswerable from 9 to 10 of 10, and multi-hop from 8 to 9 of 10.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Before and after, by category<small>Hollow is the fresh 116 run, filled is the run with the kept fix</small></p>
<p class="fig-callout"><b>+2</b>questions, both in the categories that held the pointer misses</p>
</div>
<div class="dumbbells" role="img" aria-label="Per-category pass counts, fresh run to kept fix: factual 69 to 69 of 70, trap 17 to 17, permission 13 to 13, unanswerable 9 to 10, multi-hop 8 to 9 of 10">
<div class="db-row"><span class="db-label">Factual</span><span class="db-track"><span class="db-b" style="left:98.6%"></span></span><span class="db-value">69 → 69 <small>of 70</small></span></div>
<div class="db-row"><span class="db-label">Trap</span><span class="db-track"><span class="db-b" style="left:100%"></span></span><span class="db-value">17 → 17 <small>of 17</small></span></div>
<div class="db-row"><span class="db-label">Permission</span><span class="db-track"><span class="db-b" style="left:100%"></span></span><span class="db-value">13 → 13 <small>of 13</small></span></div>
<div class="db-row is-focal"><span class="db-label">Unanswerable</span><span class="db-track"><span class="db-seg" style="left:90%;width:10%"></span><span class="db-a" style="left:90%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">9 → 10 <small>of 10</small></span></div>
<div class="db-row is-focal"><span class="db-label">Multi-hop</span><span class="db-track"><span class="db-seg" style="left:80%;width:10%"></span><span class="db-a" style="left:80%"></span><span class="db-b" style="left:90%"></span></span><span class="db-value">8 → 9 <small>of 10</small></span></div>
</div>
<figcaption>Three categories were already at or one short of their ceiling and stayed there. The gain is in the two that moved.</figcaption>
</figure>

Two misses remain, and neither has the pointer shape. `q093` is an identifier lookup where the assistant refused with the right document retrieved. `q120` is an expanded multi-hop question whose answer cited the complaint escalation document and dropped the service-level credit policy. I re-asked both, and both passed. That is run-to-run variance in a reasoning model, the same churn the August post warned about, and it is why the honest claim is a band of 116 to 120 rather than 118.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">The kept fix, one cell per question<small>118 of 120 passed</small></p>
<p class="fig-callout"><b>2</b>misses remain, and both pass when re-asked</p>
</div>
<div class="grid120" role="img" aria-label="Kept fix: 118 of 120 questions passed, 2 failed">
<div class="cg"><span class="cg-label">Factual <b>69/70</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i class="fail" title="q093"></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Trap <b>17/17</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Permission <b>13/13</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Unanswerable <b>10/10</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop <b>5/5</b></span><span class="cells"><i></i><i></i><i></i><i></i><i></i></span></div>
<div class="cg"><span class="cg-label">Multi-hop, expanded <b>4/5</b></span><span class="cells"><i></i><i></i><i></i><i></i><i class="fail" title="q120"></i></span></div>
</div>
<figcaption>The recorded run. Two misses remain: q093 and q120. Compare the six in the August post, four of which failed for every model tested.</figcaption>
</figure>

## What stayed constant

Retrieval was locked, and the numbers say so. The deterministic retrieval suite, which runs beside every live pass as a regression guard, held at exactly its August values. On the live stack, retrieval delivered the gold documents for 99.5% of questions (97.4% in the August run), with zero forbidden-document retrievals and zero degraded turns.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Retrieval, unchanged<small>Deterministic suite, identical to the August values</small></p>
</div>
<div class="stats" role="img" aria-label="Retrieval suite held at recall at 3 0.912, MRR 0.825, nDCG 0.837 and zero access-control leaks">
<div class="stat"><span class="stat-label">Recall@3</span><span class="stat-value">0.912</span><span class="stat-note">Right document in the top three</span></div>
<div class="stat"><span class="stat-label">MRR</span><span class="stat-value">0.825</span><span class="stat-note">How high the right document ranks</span></div>
<div class="stat"><span class="stat-label">nDCG</span><span class="stat-value">0.837</span><span class="stat-note">Ranking quality across the list</span></div>
<div class="stat"><span class="stat-label">Access leaks</span><span class="stat-value">0</span><span class="stat-note">Forbidden documents retrieved</span></div>
</div>
<figcaption>The same four numbers as the August post, to three decimals. Every gain in this post came from the answer layer.</figcaption>
</figure>

## What this does not show

- The 116 is a fresh run, not a measured change from the August 114. Between the two, the answer layer and retrieval went through another round of changes, recorded in the repo's execution tracker rather than in an eval report, and none of those runs is frozen. The 6 September morning run of that system scored 116; that is the number this fix was measured against. It has no frozen snapshot either; its category counts are recorded in the report's table. Only the 118 run is frozen.
- Both residual misses pass on replay, which is evidence of variance, not of a fix. Expect 116 to 120 across runs and read category totals, not single questions.
- Complete-answer latency was p50 22.9s and p95 58.9s. The report records that as comparable to August (p50 15.6s, p95 56.3s, measured beside four other model runs; a solo August segment measured p50 about 19.6s). Still above the 15s p95 production target, still an open item.
- Citation broadening is still visible: 10 of 120 grounded answers cited a document outside the gold set, one more than August's 9. Gold documents retrieved but uncited fell from 16 rows to 15.
- The grafting attempt's 4 of 4 was a smoke replay of the four failing questions, not a full run, and it has no frozen snapshot. That is deliberate: it was never a candidate for the record.
- The pointer detector was written against one synthetic corpus. Its title matching and contrast phrases are general in design, but they have only been exercised on the 65 Northwind documents.

## Evidence

The frozen findings for the 118 run are [`findings.glm-5.3-flash.json`](https://github.com/wasimjalali/useful-brain/blob/main/evals/results/2026-09-06/findings.glm-5.3-flash.json) under `evals/results/2026-09-06/`. The source report is [Pointer-triggered coverage: 116 to 118 without fitting the corpus](https://github.com/wasimjalali/useful-brain/blob/main/evals/system-evals/2026-09-06-northwind-pointer-coverage.md), the detector is [`src/lib/agent/pointer-completion.ts`](https://github.com/wasimjalali/useful-brain/blob/main/src/lib/agent/pointer-completion.ts), and the reproduce commands are in the [evals index](https://github.com/wasimjalali/useful-brain/tree/main/evals).

This is the third post about the same assistant. The first covers the repair from 72% to 95%: [From 72% to 95%: fixing an assistant that only quotes its sources](/writings/from-72-to-95-repairing-a-grounded-rag-agent/). The second ran five models through the same test and kept the cheapest of the two best: [Same score, one ninth the price: choosing a knowledge assistant's model](/writings/picking-a-model-for-a-grounded-rag-agent/).
