10 min read
From 41 hidden questions to zero: making a live Q&A filter safe to trust
A filter that sorts a live stream's Dari comments into a question queue used to hide real questions. With one strong model and a probability check before every hide, it hid none of 618 real questions and none across three recorded broadcasts.
Ṣafwa is a Chrome extension I built for a teacher who answers viewers’ questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions: greetings, prayers, thanks, the same question typed again, a question split over three messages. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order. The one thing it must never do is hide a real question from him, because a viewer who asked and was never answered simply leaves. I call that mistake a wrong hide, and every number in this post is measured against it first.
This post is about one decision: which language model should sort the comments, and what check should sit in front of it, so that wrong hides go to zero. On a set of 867 comment decisions I had labelled by hand, the model that production used at the time hid 41 of the 618 real questions when it ran alone. Gemini 3.8 Flash alone hid 6. The small “Dari-friendly” models a research report had recommended did worse than both. Adding a check that asks a second, different kind of model for a probability before any hide took Gemini to 0 of 618. Adding a second language model instead removed none of the mistakes that check already caught, and made more cards undecided. The rebuilt service, replayed end to end on three recorded broadcasts, hid 0 real questions on all three.
Real questions wrongly hidden, by setupOf 618 real questions in the same 867 labelled decisions
0real questions hidden once the probability check runs before every hide
The setup
Every comment goes through the same few decisions. In the room: is this a new question, a repeat of a question already in the queue, or a greeting? For a person who has already written: is this new message the rest of their question, an extra question, a repost, or a greeting? And a courtesy check: is this text a real question at all, or only a greeting or a prayer request? Any wrong answer that hides text is a wrong hide.
To score a model I need a test set: comments whose right answer is known. Mine is 867 decisions from eight real live sessions plus handwritten hard cases, each labelled by hand. 618 of them should keep the text visible and 249 should hide it. Every model saw the exact prompts, answer format and parser that production uses, so a model that breaks the answer format is scored as production would score it. Models that can “think” before answering had thinking switched off; Gemini 3.8 Flash cannot turn it off and ran at the lowest effort.
Four numbers matter for each setup: correct, wrong hides (of 618), should-hide left visible (of 249), and undecided. An undecided card is shown like a normal question, so it is safe but costs the teacher a glance.
Two days earlier I had run 28 models and settings, mostly on Cloudflare Workers AI, through the same prompts. Two findings from that bake-off shape this one: models with thinking left on broke the answer format so badly that the shipped backup model returned usable output on 17% of calls, and Gemma 4 26B, the production model at 94.1% correct, hid twice as many real questions (9.0% against 4.3%) on two sessions that appear in none of its prompt examples. Scores on the prompt’s own sessions flatter the incumbent.
Batch 1: a strong general model beat the small Dari picks
The trigger was a research report recommending Qwen3.5-9B first, with Gemma 3 12B named for its Kabuli and Hazaragi results. Fine-tuning one of them on my own labelled comments was on the table. First I ran the candidates alone on the 867 decisions through OpenRouter, a service that gives one API for many models: no voting, no checks, no gates. Total cost: $1.74.
Correct decisions, each model alone867 decisions through OpenRouter, production prompts and parser
98.8%for Gemini 3.8 Flash, at $1.47 of the $1.74 total
Gemini 3.8 Flash was 98.8% correct with 6 wrong hides. Half its answers came back within 1.8s and 95% within 7.9s, for $1.47 of the $1.74. Qwen3.5-9B reached 93.4% with 18 wrong hides, but broke the strict answer format on 382 of 867 calls by adding a "match": null field the production parser rejects; those calls were scored on the judgement and counted separately. Gemma 3 12B was 80.2% correct and hid 86 real questions.
The six Gemini wrong hides were the useful part. Four hid someone’s only real question because it followed a greeting from the same person. Two called a request to the teacher a greeting (“please don’t talk so fast”, “change the background”). Both patterns are exactly what a check outside the model can catch: has this person already asked a real question, and is this text free of any question or request?
So: a strong off-the-shelf model already reaches about 99% alone on this task, and fine-tuning a small Dari model is not needed now.
Combinations: a probability check, not a second model, gets to zero
For the combinations I moved Gemini 3.8 Flash and Gemma 4 26B to Google Vertex AI, where the production service would call them, and ran DeepSeek V4 Pro on Cloudflare Workers AI. Same 867 decisions, same prompts, thinking off. Alone on Vertex, Gemini scored 98.7% with 7 wrong hides (one more than through OpenRouter), DeepSeek V4 Pro 95.6% with 19, and Gemma 4 93.5% with 41.
The combinations were scored offline from the saved answers, using the production service’s own voting rules and its Jev questions and thresholds. Jev is a different kind of model, made by TypeSafe and called through OpenRouter. Instead of writing text, it takes a passage and a fixed question and returns a probability for each possible answer. That allows a hard line: a greeting may be hidden only if Jev puts the text over 0.8 on “no question and no request”, and a repeat may be folded only if Jev agrees the two texts are the same question.
| Setup | Correct | Wrong hides | Left visible | Undecided |
|---|---|---|---|---|
| Gemini alone | 98.7% | 7 | 12 | 0 |
| Gemini + Jev | 97.8% | 0 | 23 | 18 |
| Gemini + DeepSeek, both agree | 96.5% | 6 | 19 | 39 |
| Gemini + Gemma 4, both agree | 95.8% | 6 | 24 | 57 |
| Gemini + DeepSeek + Jev | 95.6% | 0 | 29 | 55 |
| Gemini, DeepSeek, Gemma 4, 2 of 3, + Jev | 96.8% | 2 | 22 | 29 |
| Gemma 4 alone (the live model) | 93.5% | 41 | 16 | 0 |
| Gemma 4 + Jev (the backup) | 95.5% | 4 | 31 | 52 |
| Gemma 4 + DeepSeek + Jev (close to the live setup) | 95.2% | 2 | 30 | 70 |
Wrong hides are out of 618 real questions, left visible out of 249 that should be hidden.
Gemini plus Jev is the only simple setup with zero wrong hides, and it beats the live-like setup on every column. The two-model rows tell the other half: the 6 wrong hides that survive “Gemini + DeepSeek, both agree” are a subset of Gemini’s own 7. Two language models agree on a wrong hide because they make the same mistake. What a second model mostly adds is disagreement, which the undecided column counts: 39 and 57 plain cards against 18.
The decision followed the numbers and the product owner’s instruction, “whatever setup gives the best result, stick to that without crowding”: Gemini 3.8 Flash as the only sorting model, plus Jev and the real-question check. The DeepSeek and GLM votes were dropped. Gemma 4 on Workers AI stays only as the backup when Vertex is unavailable, and because Gemma 4 plus Jev still hides 4 real questions (all second halves of a question it called “extra”), an extra proposed by the backup stays visible. That leaves 166 extras visible, only while Vertex is down, for 0 wrong hides.
One correction is worth recording. The first version of the combination report put Gemini plus Jev at 97.7%. A review pass found that the scorer counted a person’s repost of their own question as a hide, while the production service folds a repost onto that question and joins a partial match as a continuation. The scorer was fixed to match production; every number above is from the corrected run, and the headline stayed at 0.
End to end: three broadcasts, one wrong hide, then none
Offline scoring replays saved answers. The live service (a Cloudflare Worker, the small server the extension talks to) adds timing, a backup path and a courtesy check on the previous message, so before deploying I replayed three recorded broadcasts through the rebuilt service, uploaded as unreleased versions while production stayed on the old voting version. Broadcast A is the 2026-09-24 session: 115 comments, every outcome labelled. Broadcast B has 137 comments, 136 labelled. Broadcast C has 127, all labelled.
The rebuilt pipelineWhere each stage runs
- Gemini 3.8 Flash proposes the labelVertex AIThe only sorting model. Lowest reasoning effort, temperature 0, production answer format
- Jev must agree before any hideOpenRouterHiding a greeting needs "no question and no request" over 0.8; folding a repeat needs Jev to call the pair the same question
- Real-question check for extrasOpenRouter and Workers AIAn extra question can only be hidden if the person's earlier message was a real question. Jev scores it and Gemma 4 runs the courtesy check; either can keep the card visible
- Backup after 15 secondsWorkers AIGemma 4 26B answers when Vertex does not. Its extras stay visible and Jev still checks greetings and repeats
| Version | Broadcast | Correct | Wrong hides | Undecided | Fell back |
|---|---|---|---|---|---|
| Courtesy check on Gemini | A | 111/115 | 0 | 5 | 3 |
| Courtesy check on Gemini | B | 123/136 | 0 | 4 | 3 |
| Courtesy check on Gemini | C | 120/127 | 0 | 2 | 8 |
| Courtesy check on Gemini, A again | A | 107/115 | 1 | 7 | 19 |
| Courtesy check on Gemma 4 | A | 111/115 | 0 | 6 | 20 |
| Courtesy check on Gemma 4, repeat | A | 106/115 | 0 | 9 | 21 |
| Courtesy check on Gemma 4 | B | 123/136 | 0 | 5 | 23 |
| Courtesy check on Gemma 4 | C | 122/127 | 0 | 0 | 3 |
| Final after review, at night | A | 94/115 | 0 | 21 | 100/113 |
Real questions hidden per broadcast, end to endThe rebuilt service with the courtesy check on Gemma 4, against the old 2-of-3 voting version
The one wrong hide was one broadcast’s comment #58, a real first question. The person’s earlier message was a prayer request. Jev scored that prayer 0.78 as “no question”, just under the 0.8 line, and Gemini’s courtesy check called it a real question, so the check let the hide through. Measured directly, Gemini said “real question” on 1 of 7 answered calls on that text; Gemma 4 said “greeting” 10 of 10, as it had in the old version. So the courtesy check moved to Gemma 4 on Workers AI. It can only keep a card visible, it adds no Vertex load, and Gemini stays the only sorting model. The next two runs kept #58 visible.
Every other miss of the final version is on the safe side: a card left visible, or a second half joined to the person’s visible question. The old 2-of-3 voting version had scored 115 of 115 on broadcast A, whose comments its rules were tuned on, but 126 of 136 with one wrong hide on B and 122 of 127 on C. The new service is less decisive on A while Vertex is congested, and it has no wrong hide anywhere.
What stayed constant
The 867 labelled decisions, the production prompts, answer format and parser, and thinking off were the same in every run. Jev’s questions and thresholds are code shared between the offline scorer and the live service. The 15 second Gemini stage stayed at 15 seconds: a shorter one would fall back sooner but hand the 14 slow-but-good answers to Gemma too.
Honest caveats
- Vertex capacity, not the pipeline, now limits quality. A new project on Standard PayGo shares Google’s pool. Between 1 and 3 AM CEST, 13 of 22 direct Gemini calls got
429(the “too many requests” error), and in the replays 3 to 23 requests per broadcast (up to 18%) got no Gemini answer within the 15 second stage and were answered by Gemma. At about 3:55 AM the pre-release run fell back on 100 of 113 requests, half of them waiting 15.5s or more, and only 3 of the broadcast’s extras were hidden. Still 0 wrong hides, but most extras stayed visible. Under Google’s Standard PayGo tiers the documented fix for shared-pool429s is a paid tier. A same-day follow-up on Priority PayGo kept 0 wrong hides on all three broadcasts, but Google served only 5% of answered calls at priority, so that is not the fix yet either. - The offline scoring does not model the courtesy check on the previous message. The live pipeline is safer than the table, not less safe.
- 618 real questions is a small set for a zero. Gemini’s own count moved from 6 to 7 between OpenRouter and Vertex on identical prompts. Zero here is evidence, not a guarantee.
- Seven labelled “joined” cases end up hidden. Each continues the person’s own second question, which was itself hidden as an extra; their first real question stays visible. The old version hid all of them too.
- Gemini cannot turn thinking off. At the lowest effort, 95% of its answers arrived within 7.9s through OpenRouter and 4.8s through Vertex.
- Price. Gemini 3.8 Flash on Vertex was documented at $0.75 in and $3.75 out per million tokens through 2026-12-31, then $1.50 and $7.50, read on 2026-09-25: about $0.25 per live session, later about $0.50. The Vertex runs drew on a startup credit.
Evidence
The eval reports, the labelled dataset and every raw answer live in a private repository, because the dataset holds viewers’ comments. The numbers here come from three dated reports in its evals/results/ folder: the 2026-09-24 report on the OpenRouter batch (with openrouter-batch1-summary.json), the 2026-09-25 combinations report (with 2026-09-25-lineups.json) and the 2026-09-25 service report (with the per-broadcast replay summaries).
This is one of four posts on Ṣafwa. The 28-model comparison this decision superseded is in wrongly hidden questions 4.0% to 1.9%, the broadcast replay that fixed the confusing cards under the previous setup is in from six confusing cards to none, and the performance audit that came before all of it is in 7.8 times less CPU.