Wrongly hidden questions 4.0% to 1.9%: the test design chose the model

Twenty-eight AI models were scored on sorting a live stream's Dari comments into a question queue. The best combination cut wrongly hidden questions from 4.0% to 1.9% on the test, and four choices in how the test was built changed which model came out on top.

Ṣafwa is a Chrome extension I built for a teacher who answers viewers’ questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions: greetings, prayers, thanks, the same question typed again, a question split over three messages. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order. The one thing it must never do is hide a real question from him, because a viewer who asked and was never answered simply leaves. Every model in this comparison was measured against that mistake first.

This post is about one thing: how the test itself was designed, and how each design choice changed which model looked best. The candidates were 17 language models on Cloudflare’s Workers AI, 10 of them also run with their reasoning mode switched off, for 27 configurations, plus Jev, a different kind of model that answers with a probability instead of text. All were scored against 1,002 comment decisions I had labelled by hand, through the production server code, prompts and answer parser, 33,828 calls in all. The best combination cut wrongly hidden questions from the production setup’s 4.0% to 1.9% on this test. But four design choices each moved the ranking. Reasoning mode on broke the answer format on nine of the ten models tried both ways; the production fallback had been usable on 17% of calls. Two sessions kept out of every prompt example doubled the current model’s rate from 4.3% to 9.0%. Reporting the slowest one call in twenty instead of the typical one turned the safest model into an unusable one, at 21.7 s. And scoring a probability model on cut-offs gave a safety check with zero wrongly hidden questions. The largest remaining error was a rule the prompt stated, not a model.

Real questions wrongly hidden, on the prompt's own sessions and on two unseen sessionsShare of real questions a model would hide or fold away, from the sessions its prompt was written from (hollow) to two it had never seen (filled). Five finalists, three runs each, 0 to 10% scale

9.0%of real questions folded by the production primary on sessions it had never seen, from 4.3%

Every finalist got worse on sessions its prompt had never seen. The incumbent got worse fastest, because several of the prompt's worked examples come from the older sessions.

What was tested

The test set is 1,002 comment decisions generated by the real pipeline from eight recorded live sessions, plus hard cases I wrote by hand: a greeting with a question inside it, a question split into three messages, paraphrases, the same topic but a different question, verse and number differences, Afghan against Iranian wording, Persian typed in Latin letters, emoji-only comments and 11 attempts to smuggle instructions to the model inside a comment. Each decision has a labelled right answer. 28 are marked ambiguous and left out of every headline number. 317 come from two sessions that appear in no prompt example, which I call the unseen sessions, and 33 overlap with the prompt’s own worked examples and are reported separately.

Four decisions are scored. In the room: is this comment a new question, a repeat of question k, or a greeting? For someone who has already written: is this a continuation of their question, an extra question, a repost or a greeting? A courtesy check: is this text a real question at all? And pairs: are these two texts the same question?

Every request went through the production server with the production prompts, parser and validation. Only the model changed. A model that broke the fixed answer format was scored the way production would score it, as an invalid answer.

The protocol had three stages. Screening ran every configuration on the same 378 decisions. The final ran the top five on every decision, three times each. The pair check ran every configuration on every pair, three times. The key metric is the share of real questions (labelled as a new question or a continuation) that a model would hide or fold away. The report calls it the false-fold rate; here I call it wrongly hidden questions.

Reasoning mode broke the answer format before the model could be judged

The production fallback at the time, the model that answers when the main one fails, was GLM 5.3 Flash with its default settings. In the screening it returned an answer in the required format on 17% of calls. With its reasoning mode switched off, meaning it answers directly instead of writing out its thinking first, it returned a valid answer on 100% of calls and finished as the third-best model. The fallback had been shipped without anyone measuring whether it answered.

It was not alone. Of the ten models run both ways, nine were unusable with reasoning on and fine with it off.

ModelValid answers, reasoning onValid answers, reasoning offAccuracy, reasoning off
Qwen3 30B A3B0%100%84.9%
Kimi K2.60%100%94.1%
Nemotron 3 120B9.2%100%88.4%
gpt-oss-20b11.6%100%85.4%
gpt-oss-120b12.2%97.0%88.9%
GLM 5.314.9%100%96.8%
DeepSeek V4 Pro15.9%99.5%95.4%
GLM 5.3 Flash17.0%100%94.6%
DeepSeek V4 Flash21.6%99.7%90.8%
Qwen 3.8 27B99.7%98.6%94.9%

Answers in the required format, reasoning mode on to offShare of screening calls that came back as a label in the fixed format, 378 decisions, 0 to 100% scale

17%of the shipped fallback's calls came back usable with its default settings

Nine of the ten models run both ways were unusable with reasoning mode on and fine with it off. One setting decided more of the table than any model choice.

A public leaderboard ranks these models by what they know. This task ranks them by whether a label comes back inside a fixed format, and a paragraph of reasoning in the wrong place is a wrong answer.

Unseen sessions doubled the current model’s error

The five finalists ran on every decision three times.

ModelAccuracyWrongly hiddenWrongly hidden, unseen sessionsTypical / slowest 1 in 20Same answer across runs
DeepSeek V4 Pro, reasoning off96.5%2.8%4.3%1.1s / 3.9s96.6%
Qwen 3.8 27B94.2%2.9%3.9%0.6s / 21.7s91.7%
GLM 5.3 Flash, reasoning off95.8%3.6%4.8%0.6s / 3.2s96.6%
GLM 5.295.9%4.1%6.6%1.6s / 8.1s97.2%
Gemma 4 26B, in production94.1%5.8%9.0%0.34s / 0.76s98.9%

On the sessions the prompt was written from, Gemma 4 wrongly hid 4.3% of real questions. On the two unseen sessions it hid 9.0%. Several of the prompt’s worked examples come from the older sessions, so the model had seen close relatives of those decisions. Every finalist got worse on the unseen split, but the current model got worse fastest, and without that split this test would have overrated it. From here on, where a number has an unseen-session version, I give both.

The slowest one call in twenty is what matters live

Qwen 3.8 27B matched the best on safety, with 2.9% wrongly hidden, and had the lowest rate of all on the unseen sessions, 3.9%. Its typical answer took 0.6 s. Its slowest one call in twenty (the p95) took 21.7 s, 2.8% of its calls timed out, and it gave the same answer across runs only 91.7% of the time. A viewer’s question that waits 20 s to be sorted has already scrolled past. Reported on the typical call, Qwen is a contender. Reported on the tail, it is out.

Answer time, typical call to slowest one in twentySeconds per call, p50 (hollow) to p95 (filled), five finalists, 0 to 25 s scale

21.7sfor the slowest one call in twenty from the model with the best safety on unseen sessions

Reported on the median, Qwen 3.8 27B is a contender. Reported on the tail, one call in twenty takes over 20 s, and a question that waits that long has scrolled past.

A model that answers with a probability made the better safety check

Before a duplicate folds away, a second model is asked “are these two the same question?”. In production that model was Llama 3.3 70B. It has no official Persian support. It wrongly said “same” for 11.1% of pairs that were different questions, about one in nine. DeepSeek V4 Pro with reasoning off got that to 1.4% while still catching every real duplicate. Jev, at a threshold of 0.7, got it to 0 of 69 while catching 76% of real duplicates on the original pairs and 100% on the unseen ones, with its slowest one call in twenty at 0.36 s.

Jev is scored differently because it answers differently. It returns a probability that the comment should be hidden, so the cut-off is a choice I make on the data rather than a label the model picked. The rule was the lowest cut-off with zero wrongly hidden questions on the original sessions, checked afterwards on the unseen ones. On the greeting check the scores leave a clean gap: real questions score at most 0.67 and greetings at least 0.88, so a cut-off of 0.8 sits in the middle and gives 0 wrongly hidden out of 790 while catching every greeting.

DecisionJev, wrongly hiddenGemma 4, wrongly hidden
Repeat of an earlier question0/501 (0/37 look-alike pairs)12/501
Greeting or question0/79010/790
Continuation or extra question0/104, catching 52% of true extras14/104

Real questions wrongly hidden per decision, Jev against Gemma 4Non-ambiguous decisions minus prompt examples. Jev at its chosen cut-off per decision

0real questions wrongly hidden by Jev on every decision, at cut-offs set on the data

Used as a veto on Gemma's hides, Jev blocked all 39 of them while keeping 32 of 32 correct repeat folds and 50 of 50 correct greeting folds. On extra questions it keeps only 79 of 138, so there it is a veto, not a detector.

Used as a veto on Gemma’s hides, so that a comment is hidden only when both agree, Jev blocked all 39 of Gemma’s wrongly hidden questions while keeping 32 of 32 correct duplicate folds, 50 of 50 correct greeting folds and 31 of 33 correct pair matches. On extra questions it kept only 79 of 138, which is why the report calls it a veto there and not a detector. It never produced a malformed answer, answered in about 0.3 s, and the entire run cost $0.14.

The biggest error was a rule in the prompt, not a model

Across the five finalists and all repetitions, the most common error was labelling the continuation of a question as an extra question: 180 errors. The second was the reverse, 110. Together that is about 290 errors on the same confusion, shared by every model. The reason was in the pipeline, not the models. It capped a question at two messages, so when a viewer sent a third message the prompt told the model a continuation was not allowed, and the model answered “extra”, which hides it. Real viewers send up to six. That was a code change, made in the feed-safety batch of the audit that ran alongside this comparison: the cap is now 6, and an extra question at the cap stays visible.

Most common error patternsErrors across the five finalists and all repetitions, by gold label and the label the model gave instead

290errors on one confusion, continuation or extra question, shared by every finalist

The top two rows are the same confusion from both sides. The prompt capped a question at two messages, so the models were answering the rule they were given.

End to end, the production setup at the time (Gemma 4, GLM 5.3 Flash as it was configured, Llama 3.3 on pairs) scored 95.4% with 4.0% wrongly hidden, 6.1% on the unseen sessions, and 92.1% of its folds correct. The best combination of language models (DeepSeek V4 Pro reasoning off, GLM 5.3 Flash reasoning off, DeepSeek on pairs) scored 97.3%, 1.9% wrongly hidden, 3.6% on the unseen sessions, and 100% of its folds correct.

The report’s recommendation was to keep Gemma 4 as the main model for its speed and consistency (0.34 s typical, 300 requests a minute), put a Jev check in front of every duplicate and greeting hide at 0.7 and 0.8, switch the fallback to GLM 5.3 Flash with reasoning off, replace Llama 3.3 with DeepSeek V4 Pro on pairs, and raise the two-message cap in code.

That recommendation was later superseded. A second round in the days that followed, run on a wider set of providers, found a single stronger model and showed that a second language model in the vote added nothing Jev had not already caught. Gemini 3.8 Flash with Jev as the veto became the model. That decision is its own post, from 41 hidden questions to zero.

What stayed the same

Every configuration saw the same prompts, the same parser and the same server validation. The same 1,002 decisions, with the same 28 ambiguous ones excluded and the same 33 overlapping ones set aside, scored every model. The finalists ran three times on every decision, and the pair check three times on every pair. Connection failures were retried until coverage was complete for every reported configuration.

What this does not show

  • The sets of look-alike pairs are small. Zero errors on 37 of them still allows a true rate up to about 7.8% at one-sided 95% confidence. Proving 2% needs at least 149 zero-failure pairs, and each new live session adds a few.
  • The unseen sessions contain no duplicates between different viewers, so how many real duplicates a model catches rests mostly on handwritten cases. Of the 33 unambiguous duplicates, only 2 come from a real session.
  • Decisions were replayed without the earlier AI answers, so effects that depend on order, such as the two-message cap, do not show in per-decision scores and needed a session-level replay.
  • 7,930 local connection failures were retried and excluded.
  • Jev’s scores are not identical on repeat. They varied by up to 0.08, while its decisions at the chosen cut-offs held in 94% to 100% of repeats. Its cut-offs were tuned on the original sessions; the unseen columns are the out-of-sample check.
  • Jev is a third-party service, reached through OpenRouter, which the extension’s privacy page has to disclose.
  • The same model family through classifier.dev scored much lower in its text-only format, and its higher tier refused most long Dari inputs on the free plan, so those rows are not comparable.
  • Viewer comments are public YouTube comments. Handles and verbatim texts are not published, here or in the reports.

Evidence

The reports are evals/results/2026-09-23-summary.md, evals/results/2026-09-23-bakeoff.md with its JSON, and evals/results/2026-09-23-decision-models.md with its JSON, in the Ṣafwa repository. The repository is private because the test set holds viewers’ comments.

This is one of four posts on Ṣafwa. The performance audit that ran alongside is in 7.8 times less CPU, the broadcast replay that found the two-message cap’s cousin in production is in from six confusing cards to none, and the model decision that superseded this report is in from 41 hidden questions to zero.