Same score, one ninth the price: choosing a knowledge assistant's model

Five language models ran the same 120-question test behind a company knowledge assistant that only quotes its sources. The cheap one tied the expensive one at 95.0% correct, and two big-name models could not finish the warm-up.

Useful Brain is a question-answering assistant for company documents. It answers only in sentences quoted from those documents, with a citation on each, and says “not enough evidence” when they do not say. It runs on Cloudflare Workers AI, which hosts a menu of language models to choose from. After the repair work that took it to 114 of 120 on its live test, one question was left open: was GLM 5.3 Flash, the model I had been using, actually the right one? It had been chosen for price and availability, not by measurement.

So I ran every plausible Cloudflare-hosted chat model through the identical assistant: search the documents with a tool, quote evidence word for word, cite every sentence, refuse when the evidence is not there. Same prompt, same search stack, same corpus generation, same decoding parameters, same 120 questions with known right answers.

GLM 5.3 Flash won on the merits. It tied its full-size sibling on the best pass rate, edged it on multi-hop questions (the ones whose answer needs two documents), and costs roughly one ninth as much. The surprise was how much the task shape mattered: two frontier-scale models never made it past a ten-question warm-up, the smoke test, because they could not drive the search-then-answer loop or answer within a usable time.

From eight candidates to five finishers

5of 8 candidates ran all 120 questions

gpt-oss-120b has no chat-completions tool schema on Workers AI. Llama 4 Scout and Kimi K2.6 ran the warm-up and were dropped.

Results

ModelPassMulti-hopRetrieved recallLatency p50 / p95Price $/M in / out
GLM 5.3 Flash (incumbent)114/120 (95.0%)9/100.97415.6s / 56.3s0.15 / 0.50
GLM 5.3 (full)114/120 (95.0%)8/100.96915.3s / 56.0s1.40 / 4.40
DeepSeek V4 Flash109/120 (90.8%)6/100.94311.2s / 35.9s0.44 / 1.32
Gemma 4 26B105/120 (87.5%)6/100.92817.9s / 56.8s0.10 / 0.30
DeepSeek V4 Pro105/120 (87.5%)6/100.93314.9s / 82.3s1.32 / 3.96

Prices are Cloudflare Workers AI list pricing as documented on 31 August 2026, not measured spend.

Pass rate against output price$ per million output tokens, log scale

95.0%at $0.50 per million output tokens, roughly one ninth of GLM 5.3

85%90%95% 0.300.50124Output price, $ per M tokens, log scale Gemma 4 26B DeepSeek V4 Flash DeepSeek V4 Pro GLM 5.3 GLM 5.3 Flash
Pass rate against output price. The two GLM models share the top row; only one of them is cheap.

Method

Fairness was the whole design. Every model ran the identical agent loop, prompt, retrieval stack, corpus generation and decoding parameters: temperature 0, pinned seed, thinking disabled on quote-extraction calls where the model schema allows it.

Model selection happened through a loopback-only evalModel override on the turn endpoint (a switch that only works from the same machine), gated by a fixed allowlist and rejected outside the local trust boundary, so the locked production selection never changed during the experiment. The harness fails closed if the backend answers with a different model than requested, and per-model checkpoints pin the exact pipeline build so results from different code states cannot mix.

All runs scored all 120 questions under each question’s principal (the person the question is asked as, who may only see some documents), with zero skips and zero forbidden-document retrievals. One turn in the DeepSeek V4 Flash run degraded to keyword-only retrieval after a vector-channel error, and the harness flags that run as not baseline-comparable on retrieval.

Each candidate first ran a ten-question smoke (factual, trap, permission, unanswerable, multi-hop and identifier lookups). Models with broken tool-calling, JSON discipline or refusal behaviour were recorded and dropped; survivors ran all 120 questions.

Who did not make the full run

  • gpt-oss-120b: no chat-completions tool schema on Workers AI, so it cannot drive the search_knowledge loop through the OpenAI-compatible adapter at all.
  • Llama 4 Scout: 5 of 10 on the smoke. It skipped the search tool and returned refusals in under a second, consistent with its legacy input schema on this platform.
  • Kimi K2.6: 6 of 10 on the smoke, with answers taking up to 131 seconds and a burst of remote-binding errors during its window. Unusable latency for this loop.

That is the quiet lesson: a leaderboard-strong model is worth nothing on an agentic task if the serving platform’s schema cannot carry the tool loop.

What separated the finishers

ModelFactual /70Trap /17Permission /13Unanswerable /10Multi-hop /10
GLM 5.3 Flash661712109
GLM 5.3671712108
DeepSeek V4 Flash641712106
Gemma 4 26B621611106
DeepSeek V4 Pro621512106

Multi-hop completeness was the sharpest discriminator. The scoring rule requires citing every gold document, and only the two GLM models stayed at or above 8 of 10 (both 5 of 5 on the expanded slice). Both DeepSeeks and Gemma repeatedly answered one hop well and dropped the second citation.

Identifier lookups (error codes, clause numbers, SKU-style names) broke the non-GLM models most often. Four questions failed for every finisher, which says part of the remaining gap belongs to the corpus and the answer discipline around it, not the model choice.

Abstention discipline held everywhere. Every finisher scored 10 of 10 on unanswerable questions and at least 11 of 13 on permission questions, with zero forbidden-document retrievals across all 600 scored turns. The host-side grounding validator deserves most of that credit; it refuses anything it cannot verify verbatim, regardless of which model wrote it.

Which question failed for which model21 questions, five finishers

4questions fail for all five finishers

QuestionGLM 5.3 FlashGLM 5.3DeepSeek V4 FlashGemma 4 26BDeepSeek V4 Pro
q073 Permission
q100 Factual
q105 Factual
q110 Factual
q028 Factual
q086 Multi-hop
q088 Multi-hop
q093 Factual
q116 Multi-hop, expanded
q117 Multi-hop, expanded
q094 Factual
q095 Factual
q120 Multi-hop, expanded
q010 Factual
q051 Trap
q057 Trap
q060 Trap
q074 Permission
q089 Multi-hop
q096 Factual
q099 Factual
Every question that failed for at least one finisher, sorted by how many models it broke. The top four rows fail for all five.

DeepSeek V4 Flash is the latency pick. Fastest by a wide margin (p95 35.9s against 56s and up) at a respectable 90.8%, with the one-degraded-turn caveat attached to its retrieval numbers. If the production 15s p95 answer budget ever becomes binding, it is the named alternative.

Complete-answer latency, p50 to p950 to 90 seconds; the thin vertical line is the 15s target

Complete-answer latency, p50 to p95, on a 0 to 90 second scale. The thin vertical line is the 15 second production target. Runs executed in parallel on one worker, so treat these as relative.

Decision

Keep GLM 5.3 Flash. Equal-best quality (it ties GLM 5.3 on the total and on the expanded multi-hop slice, and edges it on locked multi-hop, 4 of 5 against 3 of 5), lowest cost of the top tier. The full GLM 5.3 offers nothing here for nine times the price. The production selection stays locked; this report is the recorded evidence behind that choice.

Honest caveats

  • The five full runs executed in parallel against one local worker, so latency numbers include contention. A solo GLM segment measured p50 about 19.6s, so treat latency columns as relative, not absolute.
  • Reasoning models on this stack are not fully deterministic even at temperature 0 with a pinned seed. Category totals are stable; individual questions can flip.
  • Citation broadening stays visible: grounded answers citing documents outside the gold set numbered 9 (GLM Flash), 8 (GLM full), 4 (DeepSeek Flash), 6 (DeepSeek Pro) and 9 (Gemma) out of 120, recorded as citedNotExpectedCount in each frozen findings file.
  • The three smoke-only results (gpt-oss-120b, Llama 4 Scout, Kimi K2.6) are recorded in prose and the execution tracker, not as frozen snapshots.
  • Total gross Workers AI spend for the campaign stayed in the low single-digit dollars against the documented unit prices. The harness bounds cost by price table rather than metering tokens per run.

Evidence

Frozen per-model findings live in the repo under evals/results/2026-08-31/. The allowlist of accepted model ids is in src/lib/models/eval-override.ts, and the reproduce commands are in the evals index.