# From six confusing cards to none: fixing a live Q&A filter by replay

> A teacher's live-stream question filter showed cards labelled "this person's 3rd question" that were neither hidden nor joined. Replaying the recorded broadcast through the real system reproduced all six, found the cause, and nine replays later the fix scored 115 of 115 with no real question hidden.

- Author: Wasim Jalali
- Published: 2026-09-25
- URL: https://wasimjalali.com/writings/from-six-confusing-cards-to-none-fixing-a-live-qa-filter-by-replay/

Ṣafwa is a Chrome extension I built for a teacher who answers viewers' questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions. Ṣafwa reads that feed and turns it into a clean queue of questions he can answer in order: repeats fold into one, a question split over several messages is joined, greetings fold away and a viewer's extra questions are flagged. After one Thursday broadcast he told me that five or six cards had carried a label saying "this person's 3rd question" or "4th question", yet were neither folded away nor joined to anything, and had never shown the "checking" state. He could not tell what the filter wanted him to do with them.

This post is about one thing: replaying that recorded broadcast through the real system, first to find the cause and then to check each fix, until the replay was clean. The first replay of the 24 September broadcast scored 108 of 115 comments correct against outcomes I had labelled by hand, showed six cards with the confusing label, and hid one real question. The ninth replay, against the version that shipped, scored 115 of 115, showed no such cards and hid nothing that should have been visible. The comments, the labels and the AI models were the same in every run. Only the server's handling of one situation, and the way the extension displayed it, changed.

<figure class="fig is-field">
<div class="fig-top">
<p class="fig-head">Correct outcomes per replay<small>The 24 September broadcast, 115 comments, scored against outcomes labelled by hand. An undecided card counts as correct wherever a new question is; every run after the first hid 0 real questions</small></p>
<p class="fig-callout"><b>115</b>of 115 correct in run 09, with no real question hidden</p>
</div>
<div class="bars" role="img" aria-label="Correct outcomes of 115 per replay run: run 01 108, run 02 113, run 03 114, run 04 114, run 05 113, run 06 112, run 07 113, run 08 98, run 09 115">
<div class="bar-row"><span class="bar-label">Run 01<small>before the fix, 6 confusing cards</small></span><span class="bar-track"><span class="bar-fill" style="width:93.9%"></span></span><span class="bar-value">1 hidden<b>108/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 02<small>2-of-3 outcome vote</small></span><span class="bar-track"><span class="bar-fill" style="width:98.3%"></span></span><span class="bar-value">4 undecided<b>113/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 03<small>plus text-free trace</small></span><span class="bar-track"><span class="bar-fill" style="width:99.1%"></span></span><span class="bar-value">2 undecided<b>114/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 04<small>review pass 1 fixes</small></span><span class="bar-track"><span class="bar-fill" style="width:99.1%"></span></span><span class="bar-value">2 undecided<b>114/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 05<small>review pass 2 fixes</small></span><span class="bar-track"><span class="bar-fill" style="width:98.3%"></span></span><span class="bar-value">3 undecided<b>113/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 06<small>released after the first change</small></span><span class="bar-track"><span class="bar-fill" style="width:97.4%"></span></span><span class="bar-value">3 undecided<b>112/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 07<small>real-question gate</small></span><span class="bar-track"><span class="bar-fill" style="width:98.3%"></span></span><span class="bar-value">3 undecided<b>113/115</b></span></div>
<div class="bar-row"><span class="bar-label">Run 08<small>greeting rule, provider error burst</small></span><span class="bar-track"><span class="bar-fill" style="width:85.2%"></span></span><span class="bar-value">17 undecided<b>98/115</b></span></div>
<div class="bar-row is-focal"><span class="bar-label">Run 09<small>plus one retry, released</small></span><span class="bar-track"><span class="bar-fill" style="width:100%"></span></span><span class="bar-value">1 undecided<b>115/115</b></span></div>
</div>
<figcaption>Every run after the first hid no real question. What the remaining runs bought was the undecided column, and run 08 is the Workers AI error burst kept in on purpose.</figcaption>
</figure>

## What the teacher saw

The broadcast ran 78 minutes and he answered roughly 60 questions from it. Five or six of the cards showed the "Nth question of this person" label without being folded or joined. Three suspects: the way the extension reads comments off the page, a second experimental way of receiving comments, or the AI decisions themselves.

## Replaying the broadcast through the real system

The public chat replay of the stream holds 115 YouTube comments from 66 people. Roughly 85 more came through other platforms and are not public, so they are not in this test. I converted the replay into the format the extension reads and played it in broadcast order through the real code: the extension's own matching step, then a request to the production server that makes the AI decisions, then the step that applies the server's answer to the card, while watching the server's logs. The first script for this became `evals/replay-worker.js`, which now fails the run if any real question ends up hidden.

Then I labelled the acceptable outcome for every comment by hand. Some comments have more than one acceptable outcome; a repost may fold or stay visible, for instance. An undecided card looks exactly like a new question on screen, so "undecided" counts as correct wherever "new question" is. Every later server change was uploaded as an unreleased version and replayed through that version before anything was released.

## The cause was a default

Not the page reading and not the experimental path. Production has the experimental path switched off, and the replay reproduced the bug with no page at all. It produced exactly six such cards.

In each of the six, the server could not settle its answer and fell back to answering "new question". Two of its models disagreed, or the confirming model timed out at 15 s, or a check rejected the proposal. The extension took "new question" on a card it had provisionally marked as an extra as a settled, visible extra, and the panel labelled every visible extra with a definitive ordinal. The "checking" chip did show, but only for the 0.4 to 2.5 s the AI took, which is why the teacher never saw it.

| Comment | What it was | Why the server answered "new question" |
| --- | --- | --- |
| #1 | the second half of a question, sent 0 s after the first | Gemma said continuation, DeepSeek said repeat |
| #65 | a real third question | the confirming model timed out |
| #85 | a repost of the person's own question, one clause shorter | the pair check said different, Jev scored 0.19 |
| #100 | a real extra question | Gemma pointed at an unrelated earlier question; the pair check rightly rejected it |
| #105 | a prayer request | Gemma said greeting, Jev scored 0.52 against a 0.8 cut-off |
| #107 | a wish to be beside the teacher on Judgment Day | Gemma said continuation, DeepSeek said greeting |

The replay found one more thing the teacher had not reported. Comment #58, a real first question, was hidden, because the person's earlier message, a plea to pray for a problem, had counted as their question, so the real one was treated as an extra.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">The same-person path after the fix<small>Where the two changes landed, from the first model's proposal to the card on screen</small></p>
</div>
<ol class="flow" aria-label="The same-person path after the fix: Gemma 4 proposes a label unchanged; DeepSeek V4 Pro and GLM 5.3 Flash confirm in parallel with a two-of-three outcome vote and one retry; a new real-question gate checks the person's earlier question with Jev and the courtesy prompt; the Worker answers with a reason code and trace; the client renders undecided cards plain with no badge">
<li class="is-locked"><span class="flow-name">Gemma 4 proposes a label</span><span class="flow-tag">unchanged</span><span class="flow-note">The same room, same-person and courtesy prompts as before. This stage was never the problem</span></li>
<li><span class="flow-name">DeepSeek V4 Pro and GLM 5.3 Flash confirm in parallel</span><span class="flow-tag">changed</span><span class="flow-note">Two of three must agree on the outcome class, hide or join, not on the exact label. A confirming model that fails fast is asked once more</span></li>
<li><span class="flow-name">Real-question gate</span><span class="flow-tag">new</span><span class="flow-note">Jev and the courtesy prompt read the person's earlier question block, which the extension now sends as <code>questionText</code>. An extra can only hide if that block was a real question, not a greeting or a prayer request</span></li>
<li><span class="flow-name">The server answers</span><span class="flow-tag">changed</span><span class="flow-note">A verdict it could not settle carries an <code>unconfirmed</code> reason code, a request id and a text-free trace of votes and scores, instead of a bare "new question"</span></li>
<li><span class="flow-name">The extension shows the card</span><span class="flow-tag">changed</span><span class="flow-note">An undecided card looks like any other question. The "Nth question of this person" badge is gone: a card is visible or folded</span></li>
</ol>
<figcaption>Only the first stage is untouched. Every later stage can now leave a card visible without labelling it, which is what the six bad cards needed.</figcaption>
</figure>

## Three consultants, one verdict

Before writing a fix I gave the evidence to three independent AI models acting as consultants: GPT-6 Astra at its highest effort in the Codex CLI, SWE-2 Max in the Devin CLI, and Fable 5.1 at high effort. All three confirmed the cause from the replay and the logs. Each added something the fix kept. Astra's point was that "not a greeting" is not positive evidence that a person has asked a question. SWE-2's was to vote on the outcome (hide or join) rather than on the exact label, and to run the confirming models in parallel. Fable's was to show undecided cards plain, with no label at all. Having read that, I removed the numbered label altogether.

## The fix

Two pull requests changed the server and the extension.

A three-model vote on the outcome. Gemma 4 proposes. DeepSeek V4 Pro and GLM 5.3 Flash answer in parallel, and two of the three must agree on the outcome, not the exact label: hide (extra question, greeting, repeat) or join (continuation, or a repost of the person's own question).

An "already asked a real question" check. Before an extra question can hide, Jev, a model that answers with a probability and counts prayer requests as not a question, and a courtesy prompt both look at the person's earlier question. The extension now sends that earlier question along as `questionText`.

A retry. A confirming model that fails fast is asked once more.

No numbered label. A card is either visible or folded, and an undecided card looks like any other question.

Diagnostics. A report the teacher can save from settings that holds no comment text, a reason code and a request id on every server answer, a text-free trace of votes and scores, and server logs switched on.

Two security fixes came out of review: a same-person request now counts double against the per-user request limit, and the privacy policy names the earlier question the extension sends. The change went through five SWE-2 review passes and two Opus passes, and every critical and high finding was fixed before release.

## Nine replays

| Run | Server under test | Correct | Real questions hidden | Confusing cards | Undecided |
| --- | --- | --- | --- | --- | --- |
| 01 | production before the fix | 108 | 1 | 6 | 0 |
| 02 | 2-of-3 vote | 113 | 0 | 0 | 4 |
| 03 | plus text-free trace | 114 | 0 | 0 | 2 |
| 04 | review pass 1 fixes | 114 | 0 | 0 | 2 |
| 05 | review pass 2 fixes | 113 | 0 | 0 | 3 |
| 06 | production after the first PR | 112 | 0 | 0 | 3 |
| 07 | real-question check | 113 | 0 | 0 | 3 |
| 08 | greeting rule change, during a Workers AI error burst | 98 | 0 | 0 | 17 |
| 09 | plus confirmer retry, released | 115 | 0 | 0 | 1 |

The first run after the vote already removed every confusing card and every wrongly hidden question. What the remaining runs bought was the undecided column: cards that were safe but cost the teacher a glance. The single undecided card in run 09 is a correct plain card.

Run 08 is kept on purpose. For a stretch, DeepSeek and GLM on Workers AI both returned errors while two replays ran at once, and 17 comments came back undecided. Every one of them stayed visible, which is the failure the design wants. It is also what led to the retry rule in run 09. A run that looks like a regression and turns out to be the safety net working is worth more in the table than out of it.

<figure class="fig">
<div class="fig-top">
<p class="fig-head">Time to sort one comment, typical to slowest one in twenty<small>Seconds from request to server answer, p50 (hollow) to p95 (filled), per replay run, 0 to 8 s scale</small></p>
<p class="fig-callout"><b>3.5s</b>for the slowest one comment in twenty in run 09, against 3.7s for the old production version in run 01</p>
</div>
<div class="dumbbells" role="img" aria-label="Per-comment round trip p50 to p95 in seconds per run: run 01 0.8 to 3.7, run 02 1.0 to 7.8, run 03 0.9 to 6.1, run 04 0.9 to 3.1, run 05 0.7 to 6.1, run 06 0.5 to 2.8, run 07 0.7 to 2.6, run 08 0.6 to 1.9, run 09 0.6 to 3.5">
<div class="db-row"><span class="db-label">Run 01</span><span class="db-track"><span class="db-seg" style="left:10%;width:36.3%"></span><span class="db-a" style="left:10%"></span><span class="db-b" style="left:46.3%"></span></span><span class="db-value">0.8 → 3.7<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 02</span><span class="db-track"><span class="db-seg" style="left:12.5%;width:85%"></span><span class="db-a" style="left:12.5%"></span><span class="db-b" style="left:97.5%"></span></span><span class="db-value">1.0 → 7.8<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 03</span><span class="db-track"><span class="db-seg" style="left:11.3%;width:65%"></span><span class="db-a" style="left:11.3%"></span><span class="db-b" style="left:76.3%"></span></span><span class="db-value">0.9 → 6.1<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 04</span><span class="db-track"><span class="db-seg" style="left:11.3%;width:27.5%"></span><span class="db-a" style="left:11.3%"></span><span class="db-b" style="left:38.8%"></span></span><span class="db-value">0.9 → 3.1<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 05</span><span class="db-track"><span class="db-seg" style="left:8.8%;width:67.5%"></span><span class="db-a" style="left:8.8%"></span><span class="db-b" style="left:76.3%"></span></span><span class="db-value">0.7 → 6.1<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 06</span><span class="db-track"><span class="db-seg" style="left:6.3%;width:28.7%"></span><span class="db-a" style="left:6.3%"></span><span class="db-b" style="left:35%"></span></span><span class="db-value">0.5 → 2.8<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 07</span><span class="db-track"><span class="db-seg" style="left:8.8%;width:23.7%"></span><span class="db-a" style="left:8.8%"></span><span class="db-b" style="left:32.5%"></span></span><span class="db-value">0.7 → 2.6<small>s</small></span></div>
<div class="db-row"><span class="db-label">Run 08</span><span class="db-track"><span class="db-seg" style="left:7.5%;width:16.3%"></span><span class="db-a" style="left:7.5%"></span><span class="db-b" style="left:23.8%"></span></span><span class="db-value">0.6 → 1.9<small>s</small></span></div>
<div class="db-row is-focal"><span class="db-label">Run 09</span><span class="db-track"><span class="db-seg" style="left:7.5%;width:36.3%"></span><span class="db-a" style="left:7.5%"></span><span class="db-b" style="left:43.8%"></span></span><span class="db-value">0.6 → 3.5<small>s</small></span></div>
</div>
<figcaption>Computed from the frozen replay files for this post; the source report does not carry timing. Run 01 counts the 113 comments that reached the server (two exact repeats never did). Two confirming models in parallel did not lengthen the typical call, the final run stays close to the old version's tail, and run 08's short tail is the error burst answering fast with nothing.</figcaption>
</figure>

Two older broadcasts show the same pattern. Replayed through the old production server, they left 7 and 18 extra questions visible with the numbered label. The final server leaves 0 undecided on the second of them. Every new hide in those replays was read by hand, and in each case the person had already asked a real question.

Run 01 scores 108 against the final labels. My first manual count said 106. Two labels were widened after review: comment #100 may also join the person's earlier question about the same scholar, and comment #81 may also fold. A label set that changes during review has to say so, or the first row of the table is not comparable with the last.

## What stayed the same

The same 115 comments in the same order, the same label file after its two widenings, and the same models throughout: Gemma 4 proposing, DeepSeek V4 Pro and GLM 5.3 Flash confirming, Jev on the checks. The extension's matching code did not change between run 01 and run 09. The undecided column moved; the hidden column stayed at 0 from run 02 on.

## What this does not show

- The replay covers the 115 public YouTube comments. The roughly 85 comments from other platforms never reach the public replay and were not tested.
- It runs without a page, so it proves the page reading was not the cause but cannot measure anything the page does.
- The labels are mine. Two were widened after review, and the "undecided counts as new" rule is a judgement about what the teacher sees, not a property of the data.
- Run 08's error burst happened while two replays ran at once against Workers AI. The count of 17 undecided says how the system fails, not how often it will.
- Run 06 scored 112 against 114 two rows earlier with the same code released. The difference is in the undecided column, which moves with how fast the models answer on the day.
- The per-run timing figure is computed from the replay files for this post and is not in the source report.
- This three-model setup was later simplified to one model, Gemini 3.8 Flash, with Jev as the veto. That decision, and the replays that checked it on three broadcasts, are in [from 41 hidden questions to zero](/writings/from-41-hidden-questions-to-zero-making-a-live-qa-filter-safe-to-trust/).

## Evidence

The report is `evals/results/2026-09-24-live-session-and-models.md` in the Ṣafwa repository, with the nine replay files, the server logs, the three consultant briefs and answers, and the review outputs in `evals/results/2026-09-24/`. The label file is `evals/sessions/expect-<session>.json`. The repository is private because the test set holds viewers' comments.

This is one of four posts on Ṣafwa. The performance audit that came first is in [7.8 times less CPU](/writings/7-8-times-less-cpu-a-live-chat-question-filter-at-400-comments/), the model comparison that chose the three models in this setup is in [wrongly hidden questions, 4.0% to 1.9%](/writings/wrongly-hidden-questions-4-0-to-1-9-percent-the-test-design-chose-the-model/), and the decision that replaced them is in [from 41 hidden questions to zero](/writings/from-41-hidden-questions-to-zero-making-a-live-qa-filter-safe-to-trust/).
