10 min read
Wrongly hidden questions 4.0% to 1.9%: the test design chose the model
Twenty-eight AI models were scored on sorting a live stream's Dari comments into a question queue. The best combination cut wrongly hidden questions from 4.0% to 1.9% on the test, and four choices in how the test was built changed which model came out on top.
Ṣafwa is a Chrome extension I built for a teacher who answers viewers’ questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions: greetings, prayers, thanks, the same question typed again, a question split over three messages. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order. The one thing it must never do is hide a real question from him, because a viewer who asked and was never answered simply leaves. Every model in this comparison was measured against that mistake first.
This post is about one thing: how the test itself was designed, and how each design choice changed which model looked best. The candidates were 17 language models on Cloudflare’s Workers AI, 10 of them also run with their reasoning mode switched off, for 27 configurations, plus Jev, a different kind of model that answers with a probability instead of text. All were scored against 1,002 comment decisions I had labelled by hand, through the production server code, prompts and answer parser, 33,828 calls in all. The best combination cut wrongly hidden questions from the production setup’s 4.0% to 1.9% on this test. But four design choices each moved the ranking. Reasoning mode on broke the answer format on nine of the ten models tried both ways; the production fallback had been usable on 17% of calls. Two sessions kept out of every prompt example doubled the current model’s rate from 4.3% to 9.0%. Reporting the slowest one call in twenty instead of the typical one turned the safest model into an unusable one, at 21.7 s. And scoring a probability model on cut-offs gave a safety check with zero wrongly hidden questions. The largest remaining error was a rule the prompt stated, not a model.
Real questions wrongly hidden, on the prompt's own sessions and on two unseen sessionsShare of real questions a model would hide or fold away, from the sessions its prompt was written from (hollow) to two it had never seen (filled). Five finalists, three runs each, 0 to 10% scale
9.0%of real questions folded by the production primary on sessions it had never seen, from 4.3%
What was tested
The test set is 1,002 comment decisions generated by the real pipeline from eight recorded live sessions, plus hard cases I wrote by hand: a greeting with a question inside it, a question split into three messages, paraphrases, the same topic but a different question, verse and number differences, Afghan against Iranian wording, Persian typed in Latin letters, emoji-only comments and 11 attempts to smuggle instructions to the model inside a comment. Each decision has a labelled right answer. 28 are marked ambiguous and left out of every headline number. 317 come from two sessions that appear in no prompt example, which I call the unseen sessions, and 33 overlap with the prompt’s own worked examples and are reported separately.
Four decisions are scored. In the room: is this comment a new question, a repeat of question k, or a greeting? For someone who has already written: is this a continuation of their question, an extra question, a repost or a greeting? A courtesy check: is this text a real question at all? And pairs: are these two texts the same question?
Every request went through the production server with the production prompts, parser and validation. Only the model changed. A model that broke the fixed answer format was scored the way production would score it, as an invalid answer.
The protocol had three stages. Screening ran every configuration on the same 378 decisions. The final ran the top five on every decision, three times each. The pair check ran every configuration on every pair, three times. The key metric is the share of real questions (labelled as a new question or a continuation) that a model would hide or fold away. The report calls it the false-fold rate; here I call it wrongly hidden questions.
Reasoning mode broke the answer format before the model could be judged
The production fallback at the time, the model that answers when the main one fails, was GLM 5.3 Flash with its default settings. In the screening it returned an answer in the required format on 17% of calls. With its reasoning mode switched off, meaning it answers directly instead of writing out its thinking first, it returned a valid answer on 100% of calls and finished as the third-best model. The fallback had been shipped without anyone measuring whether it answered.
It was not alone. Of the ten models run both ways, nine were unusable with reasoning on and fine with it off.
| Model | Valid answers, reasoning on | Valid answers, reasoning off | Accuracy, reasoning off |
|---|---|---|---|
| Qwen3 30B A3B | 0% | 100% | 84.9% |
| Kimi K2.6 | 0% | 100% | 94.1% |
| Nemotron 3 120B | 9.2% | 100% | 88.4% |
| gpt-oss-20b | 11.6% | 100% | 85.4% |
| gpt-oss-120b | 12.2% | 97.0% | 88.9% |
| GLM 5.3 | 14.9% | 100% | 96.8% |
| DeepSeek V4 Pro | 15.9% | 99.5% | 95.4% |
| GLM 5.3 Flash | 17.0% | 100% | 94.6% |
| DeepSeek V4 Flash | 21.6% | 99.7% | 90.8% |
| Qwen 3.8 27B | 99.7% | 98.6% | 94.9% |
Answers in the required format, reasoning mode on to offShare of screening calls that came back as a label in the fixed format, 378 decisions, 0 to 100% scale
17%of the shipped fallback's calls came back usable with its default settings
A public leaderboard ranks these models by what they know. This task ranks them by whether a label comes back inside a fixed format, and a paragraph of reasoning in the wrong place is a wrong answer.
Unseen sessions doubled the current model’s error
The five finalists ran on every decision three times.
| Model | Accuracy | Wrongly hidden | Wrongly hidden, unseen sessions | Typical / slowest 1 in 20 | Same answer across runs |
|---|---|---|---|---|---|
| DeepSeek V4 Pro, reasoning off | 96.5% | 2.8% | 4.3% | 1.1s / 3.9s | 96.6% |
| Qwen 3.8 27B | 94.2% | 2.9% | 3.9% | 0.6s / 21.7s | 91.7% |
| GLM 5.3 Flash, reasoning off | 95.8% | 3.6% | 4.8% | 0.6s / 3.2s | 96.6% |
| GLM 5.2 | 95.9% | 4.1% | 6.6% | 1.6s / 8.1s | 97.2% |
| Gemma 4 26B, in production | 94.1% | 5.8% | 9.0% | 0.34s / 0.76s | 98.9% |
On the sessions the prompt was written from, Gemma 4 wrongly hid 4.3% of real questions. On the two unseen sessions it hid 9.0%. Several of the prompt’s worked examples come from the older sessions, so the model had seen close relatives of those decisions. Every finalist got worse on the unseen split, but the current model got worse fastest, and without that split this test would have overrated it. From here on, where a number has an unseen-session version, I give both.
The slowest one call in twenty is what matters live
Qwen 3.8 27B matched the best on safety, with 2.9% wrongly hidden, and had the lowest rate of all on the unseen sessions, 3.9%. Its typical answer took 0.6 s. Its slowest one call in twenty (the p95) took 21.7 s, 2.8% of its calls timed out, and it gave the same answer across runs only 91.7% of the time. A viewer’s question that waits 20 s to be sorted has already scrolled past. Reported on the typical call, Qwen is a contender. Reported on the tail, it is out.
Answer time, typical call to slowest one in twentySeconds per call, p50 (hollow) to p95 (filled), five finalists, 0 to 25 s scale
21.7sfor the slowest one call in twenty from the model with the best safety on unseen sessions
A model that answers with a probability made the better safety check
Before a duplicate folds away, a second model is asked “are these two the same question?”. In production that model was Llama 3.3 70B. It has no official Persian support. It wrongly said “same” for 11.1% of pairs that were different questions, about one in nine. DeepSeek V4 Pro with reasoning off got that to 1.4% while still catching every real duplicate. Jev, at a threshold of 0.7, got it to 0 of 69 while catching 76% of real duplicates on the original pairs and 100% on the unseen ones, with its slowest one call in twenty at 0.36 s.
Jev is scored differently because it answers differently. It returns a probability that the comment should be hidden, so the cut-off is a choice I make on the data rather than a label the model picked. The rule was the lowest cut-off with zero wrongly hidden questions on the original sessions, checked afterwards on the unseen ones. On the greeting check the scores leave a clean gap: real questions score at most 0.67 and greetings at least 0.88, so a cut-off of 0.8 sits in the middle and gives 0 wrongly hidden out of 790 while catching every greeting.
| Decision | Jev, wrongly hidden | Gemma 4, wrongly hidden |
|---|---|---|
| Repeat of an earlier question | 0/501 (0/37 look-alike pairs) | 12/501 |
| Greeting or question | 0/790 | 10/790 |
| Continuation or extra question | 0/104, catching 52% of true extras | 14/104 |
Real questions wrongly hidden per decision, Jev against Gemma 4Non-ambiguous decisions minus prompt examples. Jev at its chosen cut-off per decision
0real questions wrongly hidden by Jev on every decision, at cut-offs set on the data
Used as a veto on Gemma’s hides, so that a comment is hidden only when both agree, Jev blocked all 39 of Gemma’s wrongly hidden questions while keeping 32 of 32 correct duplicate folds, 50 of 50 correct greeting folds and 31 of 33 correct pair matches. On extra questions it kept only 79 of 138, which is why the report calls it a veto there and not a detector. It never produced a malformed answer, answered in about 0.3 s, and the entire run cost $0.14.
The biggest error was a rule in the prompt, not a model
Across the five finalists and all repetitions, the most common error was labelling the continuation of a question as an extra question: 180 errors. The second was the reverse, 110. Together that is about 290 errors on the same confusion, shared by every model. The reason was in the pipeline, not the models. It capped a question at two messages, so when a viewer sent a third message the prompt told the model a continuation was not allowed, and the model answered “extra”, which hides it. Real viewers send up to six. That was a code change, made in the feed-safety batch of the audit that ran alongside this comparison: the cap is now 6, and an extra question at the cap stays visible.
Most common error patternsErrors across the five finalists and all repetitions, by gold label and the label the model gave instead
290errors on one confusion, continuation or extra question, shared by every finalist
What the report recommended, and what happened next
End to end, the production setup at the time (Gemma 4, GLM 5.3 Flash as it was configured, Llama 3.3 on pairs) scored 95.4% with 4.0% wrongly hidden, 6.1% on the unseen sessions, and 92.1% of its folds correct. The best combination of language models (DeepSeek V4 Pro reasoning off, GLM 5.3 Flash reasoning off, DeepSeek on pairs) scored 97.3%, 1.9% wrongly hidden, 3.6% on the unseen sessions, and 100% of its folds correct.
The report’s recommendation was to keep Gemma 4 as the main model for its speed and consistency (0.34 s typical, 300 requests a minute), put a Jev check in front of every duplicate and greeting hide at 0.7 and 0.8, switch the fallback to GLM 5.3 Flash with reasoning off, replace Llama 3.3 with DeepSeek V4 Pro on pairs, and raise the two-message cap in code.
That recommendation was later superseded. A second round in the days that followed, run on a wider set of providers, found a single stronger model and showed that a second language model in the vote added nothing Jev had not already caught. Gemini 3.8 Flash with Jev as the veto became the model. That decision is its own post, from 41 hidden questions to zero.
What stayed the same
Every configuration saw the same prompts, the same parser and the same server validation. The same 1,002 decisions, with the same 28 ambiguous ones excluded and the same 33 overlapping ones set aside, scored every model. The finalists ran three times on every decision, and the pair check three times on every pair. Connection failures were retried until coverage was complete for every reported configuration.
What this does not show
- The sets of look-alike pairs are small. Zero errors on 37 of them still allows a true rate up to about 7.8% at one-sided 95% confidence. Proving 2% needs at least 149 zero-failure pairs, and each new live session adds a few.
- The unseen sessions contain no duplicates between different viewers, so how many real duplicates a model catches rests mostly on handwritten cases. Of the 33 unambiguous duplicates, only 2 come from a real session.
- Decisions were replayed without the earlier AI answers, so effects that depend on order, such as the two-message cap, do not show in per-decision scores and needed a session-level replay.
- 7,930 local connection failures were retried and excluded.
- Jev’s scores are not identical on repeat. They varied by up to 0.08, while its decisions at the chosen cut-offs held in 94% to 100% of repeats. Its cut-offs were tuned on the original sessions; the unseen columns are the out-of-sample check.
- Jev is a third-party service, reached through OpenRouter, which the extension’s privacy page has to disclose.
- The same model family through classifier.dev scored much lower in its text-only format, and its higher tier refused most long Dari inputs on the free plan, so those rows are not comparable.
- Viewer comments are public YouTube comments. Handles and verbatim texts are not published, here or in the reports.
Evidence
The reports are evals/results/2026-09-23-summary.md, evals/results/2026-09-23-bakeoff.md with its JSON, and evals/results/2026-09-23-decision-models.md with its JSON, in the Ṣafwa repository. The repository is private because the test set holds viewers’ comments.
This is one of four posts on Ṣafwa. The performance audit that ran alongside is in 7.8 times less CPU, the broadcast replay that found the two-message cap’s cousin in production is in from six confusing cards to none, and the model decision that superseded this report is in from 41 hidden questions to zero.