# Wrongly hidden questions 4.0% to 1.9%: the test design chose the model

> Twenty-eight AI models were scored on sorting a live stream's Dari comments into a question queue. The best combination cut wrongly hidden questions from 4.0% to 1.9% on the test, and four choices in how the test was built changed which model came out on top.

- Author: Wasim Jalali
- Published: 2026-09-25
- URL: https://wasimjalali.com/writings/wrongly-hidden-questions-4-0-to-1-9-percent-the-test-design-chose-the-model/

Ṣafwa is a Chrome extension I built for a teacher who answers viewers' questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions: greetings, prayers, thanks, the same question typed again, a question split over three messages. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order. The one thing it must never do is hide a real question from him, because a viewer who asked and was never answered simply leaves. Every model in this comparison was measured against that mistake first.

This post is about one thing: how the test itself was designed, and how each design choice changed which model looked best. The candidates were 17 language models on Cloudflare's Workers AI, 10 of them also run with their reasoning mode switched off, for 27 configurations, plus Jev, a different kind of model that answers with a probability instead of text. All were scored against 1,002 comment decisions I had labelled by hand, through the production server code, prompts and answer parser, 33,828 calls in all. The best combination cut wrongly hidden questions from the production setup's 4.0% to 1.9% on this test. But four design choices each moved the ranking. Reasoning mode on broke the answer format on nine of the ten models tried both ways; the production fallback had been usable on 17% of calls. Two sessions kept out of every prompt example doubled the current model's rate from 4.3% to 9.0%. Reporting the slowest one call in twenty instead of the typical one turned the safest model into an unusable one, at 21.7 s. And scoring a probability model on cut-offs gave a safety check with zero wrongly hidden questions. The largest remaining error was a rule the prompt stated, not a model.

<figure class="fig is-field">
<div class="fig-top">
<p class="fig-head">Real questions wrongly hidden, on the prompt's own sessions and on two unseen sessions<small>Share of real questions a model would hide or fold away, from the sessions its prompt was written from (hollow) to two it had never seen (filled). Five finalists, three runs each, 0 to 10% scale</small></p>
<p class="fig-callout"><b>9.0%</b>of real questions folded by the production primary on sessions it had never seen, from 4.3%</p>
</div>
<div class="dumbbells" role="img" aria-label="Share of real questions wrongly hidden, original sessions to unseen sessions: DeepSeek V4 Pro reasoning off 2.1% to 4.3%, Qwen 3.8 27B 2.4% to 3.9%, GLM 5.3 Flash reasoning off 3.0% to 4.8%, GLM 5.2 2.9% to 6.6%, Gemma 4 26B the production primary 4.3% to 9.0%">
<div class="db-row"><span class="db-label">DeepSeek V4 Pro, reasoning off</span><span class="db-track"><span class="db-seg" style="left:21%;width:22%"></span><span class="db-a" style="left:21%"></span><span class="db-b" style="left:43%"></span></span><span class="db-value">2.1 → 4.3<small>%</small></span></div>
<div class="db-row"><span class="db-label">Qwen 3.8 27B</span><span class="db-track"><span class="db-seg" style="left:24%;width:15%"></span><span class="db-a" style="left:24%"></span><span class="db-b" style="left:39%"></span></span><span class="db-value">2.4 → 3.9<small>%</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.3 Flash, reasoning off</span><span class="db-track"><span class="db-seg" style="left:30%;width:18%"></span><span class="db-a" style="left:30%"></span><span class="db-b" style="left:48%"></span></span><span class="db-value">3.0 → 4.8<small>%</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.2</span><span class="db-track"><span class="db-seg" style="left:29%;width:37%"></span><span class="db-a" style="left:29%"></span><span class="db-b" style="left:66%"></span></span><span class="db-value">2.9 → 6.6<small>%</small></span></div>
<div class="db-row is-focal"><span class="db-label">Gemma 4 26B, in production</span><span class="db-track"><span class="db-seg" style="left:43%;width:47%"></span><span class="db-a" style="left:43%"></span><span class="db-b" style="left:90%"></span></span><span class="db-value">4.3 → 9.0<small>%</small></span></div>
</div>
<figcaption>Every finalist got worse on sessions its prompt had never seen. The incumbent got worse fastest, because several of the prompt's worked examples come from the older sessions.</figcaption>
</figure>

## What was tested

The test set is 1,002 comment decisions generated by the real pipeline from eight recorded live sessions, plus hard cases I wrote by hand: a greeting with a question inside it, a question split into three messages, paraphrases, the same topic but a different question, verse and number differences, Afghan against Iranian wording, Persian typed in Latin letters, emoji-only comments and 11 attempts to smuggle instructions to the model inside a comment. Each decision has a labelled right answer. 28 are marked ambiguous and left out of every headline number. 317 come from two sessions that appear in no prompt example, which I call the unseen sessions, and 33 overlap with the prompt's own worked examples and are reported separately.

Four decisions are scored. In the room: is this comment a new question, a repeat of question k, or a greeting? For someone who has already written: is this a continuation of their question, an extra question, a repost or a greeting? A courtesy check: is this text a real question at all? And pairs: are these two texts the same question?

Every request went through the production server with the production prompts, parser and validation. Only the model changed. A model that broke the fixed answer format was scored the way production would score it, as an invalid answer.

The protocol had three stages. Screening ran every configuration on the same 378 decisions. The final ran the top five on every decision, three times each. The pair check ran every configuration on every pair, three times. The key metric is the share of real questions (labelled as a new question or a continuation) that a model would hide or fold away. The report calls it the false-fold rate; here I call it wrongly hidden questions.

## Reasoning mode broke the answer format before the model could be judged

The production fallback at the time, the model that answers when the main one fails, was GLM 5.3 Flash with its default settings. In the screening it returned an answer in the required format on 17% of calls. With its reasoning mode switched off, meaning it answers directly instead of writing out its thinking first, it returned a valid answer on 100% of calls and finished as the third-best model. The fallback had been shipped without anyone measuring whether it answered.

It was not alone. Of the ten models run both ways, nine were unusable with reasoning on and fine with it off.

| Model | Valid answers, reasoning on | Valid answers, reasoning off | Accuracy, reasoning off |
| --- | --- | --- | --- |
| Qwen3 30B A3B | 0% | 100% | 84.9% |
| Kimi K2.6 | 0% | 100% | 94.1% |
| Nemotron 3 120B | 9.2% | 100% | 88.4% |
| gpt-oss-20b | 11.6% | 100% | 85.4% |
| gpt-oss-120b | 12.2% | 97.0% | 88.9% |
| GLM 5.3 | 14.9% | 100% | 96.8% |
| DeepSeek V4 Pro | 15.9% | 99.5% | 95.4% |
| GLM 5.3 Flash | 17.0% | 100% | 94.6% |
| DeepSeek V4 Flash | 21.6% | 99.7% | 90.8% |
| Qwen 3.8 27B | 99.7% | 98.6% | 94.9% |

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Answers in the required format, reasoning mode on to off<small>Share of screening calls that came back as a label in the fixed format, 378 decisions, 0 to 100% scale</small></p>
<p class="fig-callout"><b>17%</b>of the shipped fallback's calls came back usable with its default settings</p>
</div>
<div class="dumbbells" role="img" aria-label="Answers in the required format, reasoning on to reasoning off: Qwen3 30B 0% to 100%, Kimi K2.6 0% to 100%, Nemotron 3 120B 9.2% to 100%, gpt-oss-20b 11.6% to 100%, gpt-oss-120b 12.2% to 97.0%, GLM 5.3 14.9% to 100%, DeepSeek V4 Pro 15.9% to 99.5%, GLM 5.3 Flash 17.0% to 100%, DeepSeek V4 Flash 21.6% to 99.7%, Qwen 3.8 27B 99.7% to 98.6%">
<div class="db-row"><span class="db-label">Qwen3 30B A3B</span><span class="db-track"><span class="db-seg" style="left:0%;width:100%"></span><span class="db-a" style="left:0%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">0 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">Kimi K2.6</span><span class="db-track"><span class="db-seg" style="left:0%;width:100%"></span><span class="db-a" style="left:0%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">0 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">Nemotron 3 120B</span><span class="db-track"><span class="db-seg" style="left:9.2%;width:90.8%"></span><span class="db-a" style="left:9.2%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">9.2 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">gpt-oss-20b</span><span class="db-track"><span class="db-seg" style="left:11.6%;width:88.4%"></span><span class="db-a" style="left:11.6%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">11.6 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">gpt-oss-120b</span><span class="db-track"><span class="db-seg" style="left:12.2%;width:84.8%"></span><span class="db-a" style="left:12.2%"></span><span class="db-b" style="left:97%"></span></span><span class="db-value">12.2 → 97.0<small>%</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.3</span><span class="db-track"><span class="db-seg" style="left:14.9%;width:85.1%"></span><span class="db-a" style="left:14.9%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">14.9 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">DeepSeek V4 Pro</span><span class="db-track"><span class="db-seg" style="left:15.9%;width:83.6%"></span><span class="db-a" style="left:15.9%"></span><span class="db-b" style="left:99.5%"></span></span><span class="db-value">15.9 → 99.5<small>%</small></span></div>
<div class="db-row is-focal"><span class="db-label">GLM 5.3 Flash, the fallback</span><span class="db-track"><span class="db-seg" style="left:17%;width:83%"></span><span class="db-a" style="left:17%"></span><span class="db-b" style="left:100%"></span></span><span class="db-value">17.0 → 100<small>%</small></span></div>
<div class="db-row"><span class="db-label">DeepSeek V4 Flash</span><span class="db-track"><span class="db-seg" style="left:21.6%;width:78.1%"></span><span class="db-a" style="left:21.6%"></span><span class="db-b" style="left:99.7%"></span></span><span class="db-value">21.6 → 99.7<small>%</small></span></div>
<div class="db-row"><span class="db-label">Qwen 3.8 27B</span><span class="db-track"><span class="db-seg" style="left:98.6%;width:1.1%"></span><span class="db-a" style="left:99.7%"></span><span class="db-b" style="left:98.6%"></span></span><span class="db-value">99.7 → 98.6<small>%</small></span></div>
</div>
<figcaption>Nine of the ten models run both ways were unusable with reasoning mode on and fine with it off. One setting decided more of the table than any model choice.</figcaption>
</figure>

A public leaderboard ranks these models by what they know. This task ranks them by whether a label comes back inside a fixed format, and a paragraph of reasoning in the wrong place is a wrong answer.

## Unseen sessions doubled the current model's error

The five finalists ran on every decision three times.

| Model | Accuracy | Wrongly hidden | Wrongly hidden, unseen sessions | Typical / slowest 1 in 20 | Same answer across runs |
| --- | --- | --- | --- | --- | --- |
| DeepSeek V4 Pro, reasoning off | 96.5% | 2.8% | 4.3% | 1.1s / 3.9s | 96.6% |
| Qwen 3.8 27B | 94.2% | 2.9% | 3.9% | 0.6s / 21.7s | 91.7% |
| GLM 5.3 Flash, reasoning off | 95.8% | 3.6% | 4.8% | 0.6s / 3.2s | 96.6% |
| GLM 5.2 | 95.9% | 4.1% | 6.6% | 1.6s / 8.1s | 97.2% |
| Gemma 4 26B, in production | 94.1% | 5.8% | 9.0% | 0.34s / 0.76s | 98.9% |

On the sessions the prompt was written from, Gemma 4 wrongly hid 4.3% of real questions. On the two unseen sessions it hid 9.0%. Several of the prompt's worked examples come from the older sessions, so the model had seen close relatives of those decisions. Every finalist got worse on the unseen split, but the current model got worse fastest, and without that split this test would have overrated it. From here on, where a number has an unseen-session version, I give both.

## The slowest one call in twenty is what matters live

Qwen 3.8 27B matched the best on safety, with 2.9% wrongly hidden, and had the lowest rate of all on the unseen sessions, 3.9%. Its typical answer took 0.6 s. Its slowest one call in twenty (the p95) took 21.7 s, 2.8% of its calls timed out, and it gave the same answer across runs only 91.7% of the time. A viewer's question that waits 20 s to be sorted has already scrolled past. Reported on the typical call, Qwen is a contender. Reported on the tail, it is out.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Answer time, typical call to slowest one in twenty<small>Seconds per call, p50 (hollow) to p95 (filled), five finalists, 0 to 25 s scale</small></p>
<p class="fig-callout"><b>21.7s</b>for the slowest one call in twenty from the model with the best safety on unseen sessions</p>
</div>
<div class="dumbbells" role="img" aria-label="Latency p50 to p95 in seconds: Gemma 4 26B 0.34 to 0.76, GLM 5.3 Flash reasoning off 0.6 to 3.2, DeepSeek V4 Pro reasoning off 1.1 to 3.9, GLM 5.2 1.6 to 8.1, Qwen 3.8 27B 0.6 to 21.7">
<div class="db-row"><span class="db-label">Gemma 4 26B</span><span class="db-track"><span class="db-seg" style="left:1.4%;width:1.6%"></span><span class="db-a" style="left:1.4%"></span><span class="db-b" style="left:3%"></span></span><span class="db-value">0.34 → 0.76<small>s</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.3 Flash, reasoning off</span><span class="db-track"><span class="db-seg" style="left:2.4%;width:10.4%"></span><span class="db-a" style="left:2.4%"></span><span class="db-b" style="left:12.8%"></span></span><span class="db-value">0.6 → 3.2<small>s</small></span></div>
<div class="db-row"><span class="db-label">DeepSeek V4 Pro, reasoning off</span><span class="db-track"><span class="db-seg" style="left:4.4%;width:11.2%"></span><span class="db-a" style="left:4.4%"></span><span class="db-b" style="left:15.6%"></span></span><span class="db-value">1.1 → 3.9<small>s</small></span></div>
<div class="db-row"><span class="db-label">GLM 5.2</span><span class="db-track"><span class="db-seg" style="left:6.4%;width:26%"></span><span class="db-a" style="left:6.4%"></span><span class="db-b" style="left:32.4%"></span></span><span class="db-value">1.6 → 8.1<small>s</small></span></div>
<div class="db-row is-focal"><span class="db-label">Qwen 3.8 27B</span><span class="db-track"><span class="db-seg" style="left:2.4%;width:84.4%"></span><span class="db-a" style="left:2.4%"></span><span class="db-b" style="left:86.8%"></span></span><span class="db-value">0.6 → 21.7<small>s</small></span></div>
</div>
<figcaption>Reported on the median, Qwen 3.8 27B is a contender. Reported on the tail, one call in twenty takes over 20 s, and a question that waits that long has scrolled past.</figcaption>
</figure>

## A model that answers with a probability made the better safety check

Before a duplicate folds away, a second model is asked "are these two the same question?". In production that model was Llama 3.3 70B. It has no official Persian support. It wrongly said "same" for 11.1% of pairs that were different questions, about one in nine. DeepSeek V4 Pro with reasoning off got that to 1.4% while still catching every real duplicate. Jev, at a threshold of 0.7, got it to 0 of 69 while catching 76% of real duplicates on the original pairs and 100% on the unseen ones, with its slowest one call in twenty at 0.36 s.

Jev is scored differently because it answers differently. It returns a probability that the comment should be hidden, so the cut-off is a choice I make on the data rather than a label the model picked. The rule was the lowest cut-off with zero wrongly hidden questions on the original sessions, checked afterwards on the unseen ones. On the greeting check the scores leave a clean gap: real questions score at most 0.67 and greetings at least 0.88, so a cut-off of 0.8 sits in the middle and gives 0 wrongly hidden out of 790 while catching every greeting.

| Decision | Jev, wrongly hidden | Gemma 4, wrongly hidden |
| --- | --- | --- |
| Repeat of an earlier question | 0/501 (0/37 look-alike pairs) | 12/501 |
| Greeting or question | 0/790 | 10/790 |
| Continuation or extra question | 0/104, catching 52% of true extras | 14/104 |

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Real questions wrongly hidden per decision, Jev against Gemma 4<small>Non-ambiguous decisions minus prompt examples. Jev at its chosen cut-off per decision</small></p>
<p class="fig-callout"><b>0</b>real questions wrongly hidden by Jev on every decision, at cut-offs set on the data</p>
</div>
<div class="bars" role="img" aria-label="Real questions wrongly hidden per decision: repeat detection Gemma 4 12 of 501 and Jev 0 of 501; greeting against question Gemma 4 10 of 790 and Jev 0 of 790; continuation against extra Gemma 4 14 of 104 and Jev 0 of 104">
<div class="bar-row"><span class="bar-label">Repeat of an earlier question<small>Gemma 4</small></span><span class="bar-track"><span class="bar-fill" style="width:85.7%"></span></span><span class="bar-value"><b>12/501</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Repeat of an earlier question<small>Jev at 0.7</small></span><span class="bar-track"><span class="bar-fill" style="width:0%"></span></span><span class="bar-value"><b>0/501</b></span></div>
<div class="bar-row"><span class="bar-label">Greeting or question<small>Gemma 4</small></span><span class="bar-track"><span class="bar-fill" style="width:71.4%"></span></span><span class="bar-value"><b>10/790</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Greeting or question<small>Jev at 0.8</small></span><span class="bar-track"><span class="bar-fill" style="width:0%"></span></span><span class="bar-value"><b>0/790</b></span></div>
<div class="bar-row"><span class="bar-label">Continuation or extra<small>Gemma 4</small></span><span class="bar-track"><span class="bar-fill" style="width:100%"></span></span><span class="bar-value"><b>14/104</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Continuation or extra<small>Jev at 0.95, catches 52% of extras</small></span><span class="bar-track"><span class="bar-fill" style="width:0%"></span></span><span class="bar-value"><b>0/104</b></span></div>
</div>
<figcaption>Used as a veto on Gemma's hides, Jev blocked all 39 of them while keeping 32 of 32 correct repeat folds and 50 of 50 correct greeting folds. On extra questions it keeps only 79 of 138, so there it is a veto, not a detector.</figcaption>
</figure>

Used as a veto on Gemma's hides, so that a comment is hidden only when both agree, Jev blocked all 39 of Gemma's wrongly hidden questions while keeping 32 of 32 correct duplicate folds, 50 of 50 correct greeting folds and 31 of 33 correct pair matches. On extra questions it kept only 79 of 138, which is why the report calls it a veto there and not a detector. It never produced a malformed answer, answered in about 0.3 s, and the entire run cost $0.14.

## The biggest error was a rule in the prompt, not a model

Across the five finalists and all repetitions, the most common error was labelling the continuation of a question as an extra question: 180 errors. The second was the reverse, 110. Together that is about 290 errors on the same confusion, shared by every model. The reason was in the pipeline, not the models. It capped a question at two messages, so when a viewer sent a third message the prompt told the model a continuation was not allowed, and the model answered "extra", which hides it. Real viewers send up to six. That was a code change, made in the feed-safety batch of the audit that ran alongside this comparison: the cap is now 6, and an extra question at the cap stays visible.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Most common error patterns<small>Errors across the five finalists and all repetitions, by gold label and the label the model gave instead</small></p>
<p class="fig-callout"><b>290</b>errors on one confusion, continuation or extra question, shared by every finalist</p>
</div>
<div class="bars" role="img" aria-label="Most common error patterns across the five finalists: continuation labelled extra 180, extra labelled continuation 110, new question labelled duplicate 105, new question labelled greeting 63, new question with invalid output 61">
<div class="bar-row is-focal"><span class="bar-label">Continuation labelled extra<small>all 5 models</small></span><span class="bar-track"><span class="bar-fill" style="width:100%"></span></span><span class="bar-value"><b>180</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Extra labelled continuation<small>all 5 models</small></span><span class="bar-track"><span class="bar-fill" style="width:61.1%"></span></span><span class="bar-value"><b>110</b></span></div>
<div class="bar-row"><span class="bar-label">New question labelled duplicate<small>all 5 models</small></span><span class="bar-track"><span class="bar-fill" style="width:58.3%"></span></span><span class="bar-value"><b>105</b></span></div>
<div class="bar-row"><span class="bar-label">New question labelled greeting<small>all 5 models</small></span><span class="bar-track"><span class="bar-fill" style="width:35%"></span></span><span class="bar-value"><b>63</b></span></div>
<div class="bar-row"><span class="bar-label">New question, invalid output<small>3 models</small></span><span class="bar-track"><span class="bar-fill" style="width:33.9%"></span></span><span class="bar-value"><b>61</b></span></div>
</div>
<figcaption>The top two rows are the same confusion from both sides. The prompt capped a question at two messages, so the models were answering the rule they were given.</figcaption>
</figure>

## What the report recommended, and what happened next

End to end, the production setup at the time (Gemma 4, GLM 5.3 Flash as it was configured, Llama 3.3 on pairs) scored 95.4% with 4.0% wrongly hidden, 6.1% on the unseen sessions, and 92.1% of its folds correct. The best combination of language models (DeepSeek V4 Pro reasoning off, GLM 5.3 Flash reasoning off, DeepSeek on pairs) scored 97.3%, 1.9% wrongly hidden, 3.6% on the unseen sessions, and 100% of its folds correct.

The report's recommendation was to keep Gemma 4 as the main model for its speed and consistency (0.34 s typical, 300 requests a minute), put a Jev check in front of every duplicate and greeting hide at 0.7 and 0.8, switch the fallback to GLM 5.3 Flash with reasoning off, replace Llama 3.3 with DeepSeek V4 Pro on pairs, and raise the two-message cap in code.

That recommendation was later superseded. A second round in the days that followed, run on a wider set of providers, found a single stronger model and showed that a second language model in the vote added nothing Jev had not already caught. Gemini 3.8 Flash with Jev as the veto became the model. That decision is its own post, [from 41 hidden questions to zero](/writings/from-41-hidden-questions-to-zero-making-a-live-qa-filter-safe-to-trust/).

## What stayed the same

Every configuration saw the same prompts, the same parser and the same server validation. The same 1,002 decisions, with the same 28 ambiguous ones excluded and the same 33 overlapping ones set aside, scored every model. The finalists ran three times on every decision, and the pair check three times on every pair. Connection failures were retried until coverage was complete for every reported configuration.

## What this does not show

- The sets of look-alike pairs are small. Zero errors on 37 of them still allows a true rate up to about 7.8% at one-sided 95% confidence. Proving 2% needs at least 149 zero-failure pairs, and each new live session adds a few.
- The unseen sessions contain no duplicates between different viewers, so how many real duplicates a model catches rests mostly on handwritten cases. Of the 33 unambiguous duplicates, only 2 come from a real session.
- Decisions were replayed without the earlier AI answers, so effects that depend on order, such as the two-message cap, do not show in per-decision scores and needed a session-level replay.
- 7,930 local connection failures were retried and excluded.
- Jev's scores are not identical on repeat. They varied by up to 0.08, while its decisions at the chosen cut-offs held in 94% to 100% of repeats. Its cut-offs were tuned on the original sessions; the unseen columns are the out-of-sample check.
- Jev is a third-party service, reached through OpenRouter, which the extension's privacy page has to disclose.
- The same model family through classifier.dev scored much lower in its text-only format, and its higher tier refused most long Dari inputs on the free plan, so those rows are not comparable.
- Viewer comments are public YouTube comments. Handles and verbatim texts are not published, here or in the reports.

## Evidence

The reports are `evals/results/2026-09-23-summary.md`, `evals/results/2026-09-23-bakeoff.md` with its JSON, and `evals/results/2026-09-23-decision-models.md` with its JSON, in the Ṣafwa repository. The repository is private because the test set holds viewers' comments.

This is one of four posts on Ṣafwa. The performance audit that ran alongside is in [7.8 times less CPU](/writings/7-8-times-less-cpu-a-live-chat-question-filter-at-400-comments/), the broadcast replay that found the two-message cap's cousin in production is in [from six confusing cards to none](/writings/from-six-confusing-cards-to-none-fixing-a-live-qa-filter-by-replay/), and the model decision that superseded this report is in [from 41 hidden questions to zero](/writings/from-41-hidden-questions-to-zero-making-a-live-qa-filter-safe-to-trust/).
