# Same score, one ninth the price: choosing a knowledge assistant's model

> Five language models ran the same 120-question test behind a company knowledge assistant that only quotes its sources. The cheap one tied the expensive one at 95.0% correct, and two big-name models could not finish the warm-up.

- Author: Wasim Jalali
- Published: 2026-09-01
- URL: https://wasimjalali.com/writings/picking-a-model-for-a-grounded-rag-agent/

[Useful Brain](https://github.com/wasimjalali/useful-brain) is a question-answering assistant for company documents. It answers only in sentences quoted from those documents, with a citation on each, and says "not enough evidence" when they do not say. It runs on Cloudflare Workers AI, which hosts a menu of language models to choose from. After the [repair work](/writings/from-72-to-95-repairing-a-grounded-rag-agent/) that took it to 114 of 120 on its live test, one question was left open: was GLM 5.3 Flash, the model I had been using, actually the right one? It had been chosen for price and availability, not by measurement.

So I ran every plausible Cloudflare-hosted chat model through the identical assistant: search the documents with a tool, quote evidence word for word, cite every sentence, refuse when the evidence is not there. Same prompt, same search stack, same corpus generation, same decoding parameters, same 120 questions with known right answers.

GLM 5.3 Flash won on the merits. It tied its full-size sibling on the best pass rate, edged it on multi-hop questions (the ones whose answer needs two documents), and costs roughly one ninth as much. The surprise was how much the task shape mattered: two frontier-scale models never made it past a ten-question warm-up, the smoke test, because they could not drive the search-then-answer loop or answer within a usable time.

<figure class="fig is-field">
<div class="fig-top">
<p class="fig-head">From eight candidates to five finishers</p>
<p class="fig-callout"><b>5</b>of 8 candidates ran all 120 questions</p>
</div>
<div class="bars" role="img" aria-label="Eight candidates considered, five ran the full 120 questions">
<div class="bar-row"><span class="bar-label">Candidates considered</span><span class="bar-track"><span class="bar-fill" style="width:100%"></span></span><span class="bar-value"><b>8</b></span></div>
<div class="bar-row"><span class="bar-label">Could run the search-then-answer loop</span><span class="bar-track"><span class="bar-fill" style="width:87.5%"></span></span><span class="bar-value"><b>7</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Passed the warm-up, ran all 120</span><span class="bar-track"><span class="bar-fill" style="width:62.5%"></span></span><span class="bar-value"><b>5</b></span></div>
</div>
<figcaption>gpt-oss-120b has no chat-completions tool schema on Workers AI. Llama 4 Scout and Kimi K2.6 ran the warm-up and were dropped.</figcaption>
</figure>

## Results

| Model | Pass | Multi-hop | Retrieved recall | Latency p50 / p95 | Price $/M in / out |
| --- | --- | --- | --- | --- | --- |
| **GLM 5.3 Flash** (incumbent) | **114/120 (95.0%)** | 9/10 | 0.974 | 15.6s / 56.3s | 0.15 / 0.50 |
| GLM 5.3 (full) | 114/120 (95.0%) | 8/10 | 0.969 | 15.3s / 56.0s | 1.40 / 4.40 |
| DeepSeek V4 Flash | 109/120 (90.8%) | 6/10 | 0.943 | 11.2s / 35.9s | 0.44 / 1.32 |
| Gemma 4 26B | 105/120 (87.5%) | 6/10 | 0.928 | 17.9s / 56.8s | 0.10 / 0.30 |
| DeepSeek V4 Pro | 105/120 (87.5%) | 6/10 | 0.933 | 14.9s / 82.3s | 1.32 / 3.96 |

Prices are Cloudflare Workers AI list pricing as documented on 31 August 2026, not measured spend.

<figure class="fig fig-scroll">
<div class="fig-top">
<p class="fig-head">Pass rate against output price<small>$ per million output tokens, log scale</small></p>
<p class="fig-callout"><b>95.0%</b>at $0.50 per million output tokens, roughly one ninth of GLM 5.3</p>
</div>
<svg class="scatter" viewBox="0 0 640 300" role="img" aria-label="Pass rate against output price. GLM 5.3 Flash sits top-left: 95% at 0.50 dollars per million output tokens. GLM 5.3 matches the pass rate at 4.40 dollars. DeepSeek V4 Flash 90.8% at 1.32. Gemma 4 26B 87.5% at 0.30. DeepSeek V4 Pro 87.5% at 3.96.">
<g class="grid"><line x1="60" y1="250" x2="600" y2="250"/><line x1="60" y1="150" x2="600" y2="150"/><line x1="60" y1="50" x2="600" y2="50"/></g>
<g class="tick"><text x="52" y="254" text-anchor="end">85%</text><text x="52" y="154" text-anchor="end">90%</text><text x="52" y="54" text-anchor="end">95%</text></g>
<g class="tick"><text x="124" y="274" text-anchor="middle">0.30</text><text x="206" y="274" text-anchor="middle">0.50</text><text x="316" y="274" text-anchor="middle">1</text><text x="426" y="274" text-anchor="middle">2</text><text x="536" y="274" text-anchor="middle">4</text><text x="600" y="294" text-anchor="end">Output price, $ per M tokens, log scale</text></g>
<g class="pt"><circle cx="124" cy="200" r="6"/><text x="124" y="222" text-anchor="middle">Gemma 4 26B</text></g>
<g class="pt"><circle cx="360" cy="134" r="6"/><text x="360" y="156" text-anchor="middle">DeepSeek V4 Flash</text></g>
<g class="pt"><circle cx="534" cy="200" r="6"/><text x="534" y="222" text-anchor="middle">DeepSeek V4 Pro</text></g>
<g class="pt"><circle cx="551" cy="50" r="6"/><text x="551" y="30" text-anchor="middle">GLM 5.3</text></g>
<g class="pt is-focal"><circle cx="206" cy="50" r="7"/><text x="206" y="30" text-anchor="middle">GLM 5.3 Flash</text></g>
</svg>
<figcaption>Pass rate against output price. The two GLM models share the top row; only one of them is cheap.</figcaption>
</figure>

## Method

Fairness was the whole design. Every model ran the identical agent loop, prompt, retrieval stack, corpus generation and decoding parameters: temperature 0, pinned seed, thinking disabled on quote-extraction calls where the model schema allows it.

Model selection happened through a loopback-only `evalModel` override on the turn endpoint (a switch that only works from the same machine), gated by a fixed allowlist and rejected outside the local trust boundary, so the locked production selection never changed during the experiment. The harness fails closed if the backend answers with a different model than requested, and per-model checkpoints pin the exact pipeline build so results from different code states cannot mix.

All runs scored all 120 questions under each question's principal (the person the question is asked as, who may only see some documents), with zero skips and zero forbidden-document retrievals. One turn in the DeepSeek V4 Flash run degraded to keyword-only retrieval after a vector-channel error, and the harness flags that run as not baseline-comparable on retrieval.

Each candidate first ran a ten-question smoke (factual, trap, permission, unanswerable, multi-hop and identifier lookups). Models with broken tool-calling, JSON discipline or refusal behaviour were recorded and dropped; survivors ran all 120 questions.

## Who did not make the full run

- **gpt-oss-120b**: no chat-completions tool schema on Workers AI, so it cannot drive the `search_knowledge` loop through the OpenAI-compatible adapter at all.
- **Llama 4 Scout**: 5 of 10 on the smoke. It skipped the search tool and returned refusals in under a second, consistent with its legacy input schema on this platform.
- **Kimi K2.6**: 6 of 10 on the smoke, with answers taking up to 131 seconds and a burst of remote-binding errors during its window. Unusable latency for this loop.

That is the quiet lesson: a leaderboard-strong model is worth nothing on an agentic task if the serving platform's schema cannot carry the tool loop.

## What separated the finishers

| Model | Factual /70 | Trap /17 | Permission /13 | Unanswerable /10 | Multi-hop /10 |
| --- | --- | --- | --- | --- | --- |
| GLM 5.3 Flash | 66 | 17 | 12 | 10 | 9 |
| GLM 5.3 | 67 | 17 | 12 | 10 | 8 |
| DeepSeek V4 Flash | 64 | 17 | 12 | 10 | 6 |
| Gemma 4 26B | 62 | 16 | 11 | 10 | 6 |
| DeepSeek V4 Pro | 62 | 15 | 12 | 10 | 6 |

**Multi-hop completeness** was the sharpest discriminator. The scoring rule requires citing every gold document, and only the two GLM models stayed at or above 8 of 10 (both 5 of 5 on the expanded slice). Both DeepSeeks and Gemma repeatedly answered one hop well and dropped the second citation.

**Identifier lookups** (error codes, clause numbers, SKU-style names) broke the non-GLM models most often. Four questions failed for every finisher, which says part of the remaining gap belongs to the corpus and the answer discipline around it, not the model choice.

**Abstention discipline held everywhere.** Every finisher scored 10 of 10 on unanswerable questions and at least 11 of 13 on permission questions, with zero forbidden-document retrievals across all 600 scored turns. The host-side grounding validator deserves most of that credit; it refuses anything it cannot verify verbatim, regardless of which model wrote it.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Which question failed for which model<small>21 questions, five finishers</small></p>
<p class="fig-callout"><b>4</b>questions fail for all five finishers</p>
</div>
<table class="matrix" aria-label="Which of the 21 questions failed for which model"><thead><tr><th scope="col">Question</th><th scope="col">GLM 5.3 Flash</th><th scope="col">GLM 5.3</th><th scope="col">DeepSeek V4 Flash</th><th scope="col">Gemma 4 26B</th><th scope="col">DeepSeek V4 Pro</th></tr></thead><tbody><tr><th scope="row"><code>q073</code> <span>Permission</span></th><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q100</code> <span>Factual</span></th><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q105</code> <span>Factual</span></th><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q110</code> <span>Factual</span></th><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q028</code> <span>Factual</span></th><td><i></i></td><td></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q086</code> <span>Multi-hop</span></th><td></td><td></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q088</code> <span>Multi-hop</span></th><td><i></i></td><td><i></i></td><td></td><td><i></i></td><td></td></tr><tr><th scope="row"><code>q093</code> <span>Factual</span></th><td></td><td></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q116</code> <span>Multi-hop, expanded</span></th><td></td><td></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q117</code> <span>Multi-hop, expanded</span></th><td></td><td></td><td><i></i></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q094</code> <span>Factual</span></th><td></td><td></td><td><i></i></td><td><i></i></td><td></td></tr><tr><th scope="row"><code>q095</code> <span>Factual</span></th><td></td><td></td><td></td><td><i></i></td><td><i></i></td></tr><tr><th scope="row"><code>q120</code> <span>Multi-hop, expanded</span></th><td></td><td></td><td><i></i></td><td></td><td><i></i></td></tr><tr><th scope="row"><code>q010</code> <span>Factual</span></th><td></td><td></td><td></td><td></td><td><i></i></td></tr><tr><th scope="row"><code>q051</code> <span>Trap</span></th><td></td><td></td><td></td><td></td><td><i></i></td></tr><tr><th scope="row"><code>q057</code> <span>Trap</span></th><td></td><td></td><td></td><td></td><td><i></i></td></tr><tr><th scope="row"><code>q060</code> <span>Trap</span></th><td></td><td></td><td></td><td><i></i></td><td></td></tr><tr><th scope="row"><code>q074</code> <span>Permission</span></th><td></td><td></td><td></td><td><i></i></td><td></td></tr><tr><th scope="row"><code>q089</code> <span>Multi-hop</span></th><td></td><td><i></i></td><td></td><td></td><td></td></tr><tr><th scope="row"><code>q096</code> <span>Factual</span></th><td></td><td></td><td></td><td></td><td><i></i></td></tr><tr><th scope="row"><code>q099</code> <span>Factual</span></th><td></td><td></td><td></td><td><i></i></td><td></td></tr></tbody></table>
<figcaption>Every question that failed for at least one finisher, sorted by how many models it broke. The top four rows fail for all five.</figcaption>
</figure>

**DeepSeek V4 Flash is the latency pick.** Fastest by a wide margin (p95 35.9s against 56s and up) at a respectable 90.8%, with the one-degraded-turn caveat attached to its retrieval numbers. If the production 15s p95 answer budget ever becomes binding, it is the named alternative.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Complete-answer latency, p50 to p95<small>0 to 90 seconds; the thin vertical line is the 15s target</small></p>
</div>
<div class="dumbbells has-target" role="img" aria-label="Latency p50 to p95 per model, on a 0 to 90 second scale, with the 15 second target marked. DeepSeek V4 Flash 11.2 to 35.9 seconds is fastest; DeepSeek V4 Pro 14.9 to 82.3 seconds is slowest at p95.">
<div class="db-row"><span class="db-label">GLM 5.3 Flash</span><span class="db-track"><span class="db-target" style="left:16.7%"></span><span class="db-seg" style="left:17.3%;width:45.3%"></span><span class="db-a" style="left:17.3%"></span><span class="db-b" style="left:62.6%"></span></span><span class="db-value">15.6 → 56.3<small>s</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.3</span><span class="db-track"><span class="db-target" style="left:16.7%"></span><span class="db-seg" style="left:17%;width:45.2%"></span><span class="db-a" style="left:17%"></span><span class="db-b" style="left:62.2%"></span></span><span class="db-value">15.3 → 56.0<small>s</small></span></div>
<div class="db-row"><span class="db-label">DeepSeek V4 Flash</span><span class="db-track"><span class="db-target" style="left:16.7%"></span><span class="db-seg" style="left:12.4%;width:27.5%"></span><span class="db-a" style="left:12.4%"></span><span class="db-b" style="left:39.9%"></span></span><span class="db-value">11.2 → 35.9<small>s</small></span></div>
<div class="db-row"><span class="db-label">Gemma 4 26B</span><span class="db-track"><span class="db-target" style="left:16.7%"></span><span class="db-seg" style="left:19.9%;width:43.2%"></span><span class="db-a" style="left:19.9%"></span><span class="db-b" style="left:63.1%"></span></span><span class="db-value">17.9 → 56.8<small>s</small></span></div>
<div class="db-row"><span class="db-label">DeepSeek V4 Pro</span><span class="db-track"><span class="db-target" style="left:16.7%"></span><span class="db-seg" style="left:16.6%;width:74.8%"></span><span class="db-a" style="left:16.6%"></span><span class="db-b" style="left:91.4%"></span></span><span class="db-value">14.9 → 82.3<small>s</small></span></div>
</div>
<figcaption>Complete-answer latency, p50 to p95, on a 0 to 90 second scale. The thin vertical line is the 15 second production target. Runs executed in parallel on one worker, so treat these as relative.</figcaption>
</figure>

## Decision

Keep GLM 5.3 Flash. Equal-best quality (it ties GLM 5.3 on the total and on the expanded multi-hop slice, and edges it on locked multi-hop, 4 of 5 against 3 of 5), lowest cost of the top tier. The full GLM 5.3 offers nothing here for nine times the price. The production selection stays locked; this report is the recorded evidence behind that choice.

## Honest caveats

- The five full runs executed in parallel against one local worker, so latency numbers include contention. A solo GLM segment measured p50 about 19.6s, so treat latency columns as relative, not absolute.
- Reasoning models on this stack are not fully deterministic even at temperature 0 with a pinned seed. Category totals are stable; individual questions can flip.
- Citation broadening stays visible: grounded answers citing documents outside the gold set numbered 9 (GLM Flash), 8 (GLM full), 4 (DeepSeek Flash), 6 (DeepSeek Pro) and 9 (Gemma) out of 120, recorded as `citedNotExpectedCount` in each frozen findings file.
- The three smoke-only results (gpt-oss-120b, Llama 4 Scout, Kimi K2.6) are recorded in prose and the execution tracker, not as frozen snapshots.
- Total gross Workers AI spend for the campaign stayed in the low single-digit dollars against the documented unit prices. The harness bounds cost by price table rather than metering tokens per run.

## Evidence

Frozen per-model findings live in the repo under [`evals/results/2026-08-31/`](https://github.com/wasimjalali/useful-brain/tree/main/evals/results/2026-08-31). The allowlist of accepted model ids is in [`src/lib/models/eval-override.ts`](https://github.com/wasimjalali/useful-brain/blob/main/src/lib/models/eval-override.ts), and the reproduce commands are in the [evals index](https://github.com/wasimjalali/useful-brain/tree/main/evals).
