From six confusing cards to none: fixing a live Q&A filter by replay

A teacher's live-stream question filter showed cards labelled "this person's 3rd question" that were neither hidden nor joined. Replaying the recorded broadcast through the real system reproduced all six, found the cause, and nine replays later the fix scored 115 of 115 with no real question hidden.

Ṣafwa is a Chrome extension I built for a teacher who answers viewers’ questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order: repeats fold into one, a question split over several messages is joined, greetings fold away and a viewer’s extra questions are flagged. After one Thursday broadcast he told me that five or six cards had carried a label saying “this person’s 3rd question” or “4th question”, yet were neither folded away nor joined to anything, and had never shown the “checking” state. He could not tell what the filter wanted him to do with them.

This post is about one thing: replaying that recorded broadcast through the real system, first to find the cause and then to check each fix, until the replay was clean. The first replay of the 24 September broadcast scored 108 of 115 comments correct against outcomes I had labelled by hand, showed six cards with the confusing label, and hid one real question. The ninth replay, against the version that shipped, scored 115 of 115, showed no such cards and hid nothing that should have been visible. The comments, the labels and the AI models were the same in every run. Only the server’s handling of one situation, and the way the extension displayed it, changed.

Correct outcomes per replayThe 24 September broadcast, 115 comments, scored against outcomes labelled by hand. An undecided card counts as correct wherever a new question is; every run after the first hid 0 real questions

115of 115 correct in run 09, with no real question hidden

Every run after the first hid no real question. What the remaining runs bought was the undecided column, and run 08 is the Workers AI error burst kept in on purpose.

What the teacher saw

The broadcast ran 78 minutes and he answered roughly 60 questions from it. Five or six of the cards showed the “Nth question of this person” label without being folded or joined. Three suspects: the way the extension reads comments off the page, a second experimental way of receiving comments, or the AI decisions themselves.

Replaying the broadcast through the real system

The public chat replay of the stream holds 115 YouTube comments from 66 people. Roughly 85 more came through other platforms and are not public, so they are not in this test. I converted the replay into the format the extension reads and played it in broadcast order through the real code: the extension’s own matching step, then a request to the production server that makes the AI decisions, then the step that applies the server’s answer to the card, while watching the server’s logs. The first script for this became evals/replay-worker.js, which now fails the run if any real question ends up hidden.

Then I labelled the acceptable outcome for every comment by hand. Some comments have more than one acceptable outcome; a repost may fold or stay visible, for instance. An undecided card looks exactly like a new question on screen, so “undecided” counts as correct wherever “new question” is. Every later server change was uploaded as an unreleased version and replayed through that version before anything was released.

The cause was a default

Not the page reading and not the experimental path. Production has the experimental path switched off, and the replay reproduced the bug with no page at all. It produced exactly six such cards.

In each of the six, the server could not settle its answer and fell back to answering “new question”. Two of its models disagreed, or the confirming model timed out at 15 s, or a check rejected the proposal. The extension took “new question” on a card it had provisionally marked as an extra as a settled, visible extra, and the panel labelled every visible extra with a definitive ordinal. The “checking” chip did show, but only for the 0.4 to 2.5 s the AI took, which is why the teacher never saw it.

CommentWhat it wasWhy the server answered “new question”
#1the second half of a question, sent 0 s after the firstGemma said continuation, DeepSeek said repeat
#65a real third questionthe confirming model timed out
#85a repost of the person’s own question, one clause shorterthe pair check said different, Jev scored 0.19
#100a real extra questionGemma pointed at an unrelated earlier question; the pair check rightly rejected it
#105a prayer requestGemma said greeting, Jev scored 0.52 against a 0.8 cut-off
#107a wish to be beside the teacher on Judgment DayGemma said continuation, DeepSeek said greeting

The replay found one more thing the teacher had not reported. Comment #58, a real first question, was hidden, because the person’s earlier message, a plea to pray for a problem, had counted as their question, so the real one was treated as an extra.

The same-person path after the fixWhere the two changes landed, from the first model's proposal to the card on screen

  1. Gemma 4 proposes a labelunchangedThe same room, same-person and courtesy prompts as before. This stage was never the problem
  2. DeepSeek V4 Pro and GLM 5.3 Flash confirm in parallelchangedTwo of three must agree on the outcome class, hide or join, not on the exact label. A confirming model that fails fast is asked once more
  3. Real-question gatenewJev and the courtesy prompt read the person's earlier question block, which the extension now sends as questionText. An extra can only hide if that block was a real question, not a greeting or a prayer request
  4. The server answerschangedA verdict it could not settle carries an unconfirmed reason code, a request id and a text-free trace of votes and scores, instead of a bare "new question"
  5. The extension shows the cardchangedAn undecided card looks like any other question. The "Nth question of this person" badge is gone: a card is visible or folded
Only the first stage is untouched. Every later stage can now leave a card visible without labelling it, which is what the six bad cards needed.

Three consultants, one verdict

Before writing a fix I gave the evidence to three independent AI models acting as consultants: GPT-6 Astra at its highest effort in the Codex CLI, SWE-2 Max in the Devin CLI, and Fable 5.1 at high effort. All three confirmed the cause from the replay and the logs. Each added something the fix kept. Astra’s point was that “not a greeting” is not positive evidence that a person has asked a question. SWE-2’s was to vote on the outcome (hide or join) rather than on the exact label, and to run the confirming models in parallel. Fable’s was to show undecided cards plain, with no label at all. Having read that, I removed the numbered label altogether.

The fix

Two pull requests changed the server and the extension.

A three-model vote on the outcome. Gemma 4 proposes. DeepSeek V4 Pro and GLM 5.3 Flash answer in parallel, and two of the three must agree on the outcome, not the exact label: hide (extra question, greeting, repeat) or join (continuation, or a repost of the person’s own question).

An “already asked a real question” check. Before an extra question can hide, Jev, a model that answers with a probability and counts prayer requests as not a question, and a courtesy prompt both look at the person’s earlier question. The extension now sends that earlier question along as questionText.

A retry. A confirming model that fails fast is asked once more.

No numbered label. A card is either visible or folded, and an undecided card looks like any other question.

Diagnostics. A report the teacher can save from settings that holds no comment text, a reason code and a request id on every server answer, a text-free trace of votes and scores, and server logs switched on.

Two security fixes came out of review: a same-person request now counts double against the per-user request limit, and the privacy policy names the earlier question the extension sends. The change went through five SWE-2 review passes and two Opus passes, and every critical and high finding was fixed before release.

Nine replays

RunServer under testCorrectReal questions hiddenConfusing cardsUndecided
01production before the fix108160
022-of-3 vote113004
03plus text-free trace114002
04review pass 1 fixes114002
05review pass 2 fixes113003
06production after the first PR112003
07real-question check113003
08greeting rule change, during a Workers AI error burst980017
09plus confirmer retry, released115001

The first run after the vote already removed every confusing card and every wrongly hidden question. What the remaining runs bought was the undecided column: cards that were safe but cost the teacher a glance. The single undecided card in run 09 is a correct plain card.

Run 08 is kept on purpose. For a stretch, DeepSeek and GLM on Workers AI both returned errors while two replays ran at once, and 17 comments came back undecided. Every one of them stayed visible, which is the failure the design wants. It is also what led to the retry rule in run 09. A run that looks like a regression and turns out to be the safety net working is worth more in the table than out of it.

Time to sort one comment, typical to slowest one in twentySeconds from request to server answer, p50 (hollow) to p95 (filled), per replay run, 0 to 8 s scale

3.5sfor the slowest one comment in twenty in run 09, against 3.7s for the old production version in run 01

Computed from the frozen replay files for this post; the source report does not carry timing. Run 01 counts the 113 comments that reached the server (two exact repeats never did). Two confirming models in parallel did not lengthen the typical call, the final run stays close to the old version's tail, and run 08's short tail is the error burst answering fast with nothing.

Two older broadcasts show the same pattern. Replayed through the old production server, they left 7 and 18 extra questions visible with the numbered label. The final server leaves 0 undecided on the second of them. Every new hide in those replays was read by hand, and in each case the person had already asked a real question.

Run 01 scores 108 against the final labels. My first manual count said 106. Two labels were widened after review: comment #100 may also join the person’s earlier question about the same scholar, and comment #81 may also fold. A label set that changes during review has to say so, or the first row of the table is not comparable with the last.

What stayed the same

The same 115 comments in the same order, the same label file after its two widenings, and the same models throughout: Gemma 4 proposing, DeepSeek V4 Pro and GLM 5.3 Flash confirming, Jev on the checks. The extension’s matching code did not change between run 01 and run 09. The undecided column moved; the hidden column stayed at 0 from run 02 on.

What this does not show

  • The replay covers the 115 public YouTube comments. The roughly 85 comments from other platforms never reach the public replay and were not tested.
  • It runs without a page, so it proves the page reading was not the cause but cannot measure anything the page does.
  • The labels are mine. Two were widened after review, and the “undecided counts as new” rule is a judgement about what the teacher sees, not a property of the data.
  • Run 08’s error burst happened while two replays ran at once against Workers AI. The count of 17 undecided says how the system fails, not how often it will.
  • Run 06 scored 112 against 114 two rows earlier with the same code released. The difference is in the undecided column, which moves with how fast the models answer on the day.
  • The per-run timing figure is computed from the replay files for this post and is not in the source report.
  • This three-model setup was later simplified to one model, Gemini 3.8 Flash, with Jev as the veto. That decision, and the replays that checked it on three broadcasts, are in from 41 hidden questions to zero.

Evidence

The report is evals/results/2026-09-24-live-session-and-models.md in the Ṣafwa repository, with the nine replay files, the server logs, the three consultant briefs and answers, and the review outputs in evals/results/2026-09-24/. The label file is evals/sessions/expect-<session>.json. The repository is private because the test set holds viewers’ comments.

This is one of four posts on Ṣafwa. The performance audit that came first is in 7.8 times less CPU, the model comparison that chose the three models in this setup is in wrongly hidden questions, 4.0% to 1.9%, and the decision that replaced them is in from 41 hidden questions to zero.