5 min read
Same score, one ninth the price: choosing a knowledge assistant's model
Five language models ran the same 120-question test behind a company knowledge assistant that only quotes its sources. The cheap one tied the expensive one at 95.0% correct, and two big-name models could not finish the warm-up.
Useful Brain is a question-answering assistant for company documents. It answers only in sentences quoted from those documents, with a citation on each, and says “not enough evidence” when they do not say. It runs on Cloudflare Workers AI, which hosts a menu of language models to choose from. After the repair work that took it to 114 of 120 on its live test, one question was left open: was GLM 5.3 Flash, the model I had been using, actually the right one? It had been chosen for price and availability, not by measurement.
So I ran every plausible Cloudflare-hosted chat model through the identical assistant: search the documents with a tool, quote evidence word for word, cite every sentence, refuse when the evidence is not there. Same prompt, same search stack, same corpus generation, same decoding parameters, same 120 questions with known right answers.
GLM 5.3 Flash won on the merits. It tied its full-size sibling on the best pass rate, edged it on multi-hop questions (the ones whose answer needs two documents), and costs roughly one ninth as much. The surprise was how much the task shape mattered: two frontier-scale models never made it past a ten-question warm-up, the smoke test, because they could not drive the search-then-answer loop or answer within a usable time.
From eight candidates to five finishers
5of 8 candidates ran all 120 questions
Results
| Model | Pass | Multi-hop | Retrieved recall | Latency p50 / p95 | Price $/M in / out |
|---|---|---|---|---|---|
| GLM 5.3 Flash (incumbent) | 114/120 (95.0%) | 9/10 | 0.974 | 15.6s / 56.3s | 0.15 / 0.50 |
| GLM 5.3 (full) | 114/120 (95.0%) | 8/10 | 0.969 | 15.3s / 56.0s | 1.40 / 4.40 |
| DeepSeek V4 Flash | 109/120 (90.8%) | 6/10 | 0.943 | 11.2s / 35.9s | 0.44 / 1.32 |
| Gemma 4 26B | 105/120 (87.5%) | 6/10 | 0.928 | 17.9s / 56.8s | 0.10 / 0.30 |
| DeepSeek V4 Pro | 105/120 (87.5%) | 6/10 | 0.933 | 14.9s / 82.3s | 1.32 / 3.96 |
Prices are Cloudflare Workers AI list pricing as documented on 31 August 2026, not measured spend.
Pass rate against output price$ per million output tokens, log scale
95.0%at $0.50 per million output tokens, roughly one ninth of GLM 5.3
Method
Fairness was the whole design. Every model ran the identical agent loop, prompt, retrieval stack, corpus generation and decoding parameters: temperature 0, pinned seed, thinking disabled on quote-extraction calls where the model schema allows it.
Model selection happened through a loopback-only evalModel override on the turn endpoint (a switch that only works from the same machine), gated by a fixed allowlist and rejected outside the local trust boundary, so the locked production selection never changed during the experiment. The harness fails closed if the backend answers with a different model than requested, and per-model checkpoints pin the exact pipeline build so results from different code states cannot mix.
All runs scored all 120 questions under each question’s principal (the person the question is asked as, who may only see some documents), with zero skips and zero forbidden-document retrievals. One turn in the DeepSeek V4 Flash run degraded to keyword-only retrieval after a vector-channel error, and the harness flags that run as not baseline-comparable on retrieval.
Each candidate first ran a ten-question smoke (factual, trap, permission, unanswerable, multi-hop and identifier lookups). Models with broken tool-calling, JSON discipline or refusal behaviour were recorded and dropped; survivors ran all 120 questions.
Who did not make the full run
- gpt-oss-120b: no chat-completions tool schema on Workers AI, so it cannot drive the
search_knowledgeloop through the OpenAI-compatible adapter at all. - Llama 4 Scout: 5 of 10 on the smoke. It skipped the search tool and returned refusals in under a second, consistent with its legacy input schema on this platform.
- Kimi K2.6: 6 of 10 on the smoke, with answers taking up to 131 seconds and a burst of remote-binding errors during its window. Unusable latency for this loop.
That is the quiet lesson: a leaderboard-strong model is worth nothing on an agentic task if the serving platform’s schema cannot carry the tool loop.
What separated the finishers
| Model | Factual /70 | Trap /17 | Permission /13 | Unanswerable /10 | Multi-hop /10 |
|---|---|---|---|---|---|
| GLM 5.3 Flash | 66 | 17 | 12 | 10 | 9 |
| GLM 5.3 | 67 | 17 | 12 | 10 | 8 |
| DeepSeek V4 Flash | 64 | 17 | 12 | 10 | 6 |
| Gemma 4 26B | 62 | 16 | 11 | 10 | 6 |
| DeepSeek V4 Pro | 62 | 15 | 12 | 10 | 6 |
Multi-hop completeness was the sharpest discriminator. The scoring rule requires citing every gold document, and only the two GLM models stayed at or above 8 of 10 (both 5 of 5 on the expanded slice). Both DeepSeeks and Gemma repeatedly answered one hop well and dropped the second citation.
Identifier lookups (error codes, clause numbers, SKU-style names) broke the non-GLM models most often. Four questions failed for every finisher, which says part of the remaining gap belongs to the corpus and the answer discipline around it, not the model choice.
Abstention discipline held everywhere. Every finisher scored 10 of 10 on unanswerable questions and at least 11 of 13 on permission questions, with zero forbidden-document retrievals across all 600 scored turns. The host-side grounding validator deserves most of that credit; it refuses anything it cannot verify verbatim, regardless of which model wrote it.
Which question failed for which model21 questions, five finishers
4questions fail for all five finishers
| Question | GLM 5.3 Flash | GLM 5.3 | DeepSeek V4 Flash | Gemma 4 26B | DeepSeek V4 Pro |
|---|---|---|---|---|---|
q073 Permission | |||||
q100 Factual | |||||
q105 Factual | |||||
q110 Factual | |||||
q028 Factual | |||||
q086 Multi-hop | |||||
q088 Multi-hop | |||||
q093 Factual | |||||
q116 Multi-hop, expanded | |||||
q117 Multi-hop, expanded | |||||
q094 Factual | |||||
q095 Factual | |||||
q120 Multi-hop, expanded | |||||
q010 Factual | |||||
q051 Trap | |||||
q057 Trap | |||||
q060 Trap | |||||
q074 Permission | |||||
q089 Multi-hop | |||||
q096 Factual | |||||
q099 Factual |
DeepSeek V4 Flash is the latency pick. Fastest by a wide margin (p95 35.9s against 56s and up) at a respectable 90.8%, with the one-degraded-turn caveat attached to its retrieval numbers. If the production 15s p95 answer budget ever becomes binding, it is the named alternative.
Complete-answer latency, p50 to p950 to 90 seconds; the thin vertical line is the 15s target
Decision
Keep GLM 5.3 Flash. Equal-best quality (it ties GLM 5.3 on the total and on the expanded multi-hop slice, and edges it on locked multi-hop, 4 of 5 against 3 of 5), lowest cost of the top tier. The full GLM 5.3 offers nothing here for nine times the price. The production selection stays locked; this report is the recorded evidence behind that choice.
Honest caveats
- The five full runs executed in parallel against one local worker, so latency numbers include contention. A solo GLM segment measured p50 about 19.6s, so treat latency columns as relative, not absolute.
- Reasoning models on this stack are not fully deterministic even at temperature 0 with a pinned seed. Category totals are stable; individual questions can flip.
- Citation broadening stays visible: grounded answers citing documents outside the gold set numbered 9 (GLM Flash), 8 (GLM full), 4 (DeepSeek Flash), 6 (DeepSeek Pro) and 9 (Gemma) out of 120, recorded as
citedNotExpectedCountin each frozen findings file. - The three smoke-only results (gpt-oss-120b, Llama 4 Scout, Kimi K2.6) are recorded in prose and the execution tracker, not as frozen snapshots.
- Total gross Workers AI spend for the campaign stayed in the low single-digit dollars against the documented unit prices. The harness bounds cost by price table rather than metering tokens per run.
Evidence
Frozen per-model findings live in the repo under evals/results/2026-08-31/. The allowlist of accepted model ids is in src/lib/models/eval-override.ts, and the reproduce commands are in the evals index.