7.8 times less CPU: a live-chat question filter at 400 comments

A Chrome extension that sorts a live stream's comments into a question queue was slowing the broadcast tab down as the chat grew. After four AI models audited it and seven rounds of fixes, a 400-comment session takes 7.8 times less processor time.

Ṣafwa is a Chrome extension I built for a teacher who answers viewers’ questions during live streams in Dari and Persian. Hundreds of comments scroll past in an hour, and most are not questions. Ṣafwa reads that feed and turns it into a clean queue: repeated questions fold into one, a question split over several messages is joined, greetings fold away and a viewer’s extra questions are flagged. It does all of that inside the same browser tab that runs the broadcast. The problem was that the longer a stream ran, the more of that tab’s processor time the filter took, until by a few hundred comments it was doing work the teacher could feel as stutter.

This post is about one thing: how much processor time and delay the filter costs, measured the same way before and after a performance audit by four AI models and seven rounds of fixes. At 400 comments the original spent 33,174 ms of processor time, took 151.4 ms to handle each new comment by the end of the session, kept the tab busy for 418 ms at its longest single stretch, and sent 42.7 MB of data to the side panel. After the seventh round the same session spent 4,240 ms, took 17.9 ms per comment, peaked at 45.5 ms and sent 3.7 MB. That is 7.8 times less processor time at 400 comments and 7.0 times less at 200. The matching rules, the recorded comments and the simulated AI were identical in both measurements.

Processor time per 400-comment sessionMilliseconds of CPU on the test rig: the extension's real session code, a fake page, a simulated AI answering in 3 s, Apple M1

4,240ms of processor time after batch 7, from 33,174. 7.8 times less

Batch 2 is missing because it touched only the side panel, which the rig cannot see. Three batches moved the number: the caches, the skipped replays and the fingerprint index.

How I measured it

The measurements come from a test rig: a script that runs the extension’s real session code outside the browser, against a fake page and a simulated AI that answers every request in 3 s. It feeds recorded comments in bursts and reports four things per run. Total processor time. Milliseconds per comment for the first 20 and the last 20 comments, which shows whether the cost grows. The longest single task, meaning the longest stretch the pipeline kept the tab busy without a pause; anything over 50 ms is visible as stutter. And the bytes sent to the side panel. Every run used the same recorded comments on an idle Apple M1, measured back to back.

The rig cannot see everything. It measures no page layout, no drawing and no cost of passing messages between browser processes, so anything on the panel side needed a separate check in a real browser. I come back to that in the caveats.

Before: every new comment cost more than the last

Each time the AI answered about one comment, the extension replayed the whole session from the start to work out what the answer changed. Each replay re-stripped the polite titles Dari speakers put before a name (the honorifics) from every stored comment and re-split every candidate into words for the near-duplicate check. So the cost of one answer grew with the number of comments already seen, and the number of answers grew with the comments too.

CommentsprocessComment callsms per comment, first 20 / last 20Longest taskData to panel
501,4304.0 / 12.012.6 ms0.8 MB
1005,3023.8 / 27.528.6 ms3.0 MB
20020,3924.7 / 70.071.8 ms11.3 MB
40076,8384.0 / 138.3165.8 ms42.7 MB

The first 20 comments cost about 4 ms each whatever the session length. The last 20 cost 138.3 ms each at 400 comments, roughly 35 times more, and single tasks had crossed the 50 ms stutter line by comment 200. The teacher’s laptop is slower than the machine I measured on, so it crosses that line sooner.

Cost per comment, start to end of a sessionms per comment for the first 20 (hollow) to the last 20 (filled), on a 0 to 150 ms scale. The line marks 50 ms, where one task becomes visible jank

The first comments cost the same at every session length. The last ones cost 35 times more at 400 comments, because every AI answer replayed everything before it.

Four auditors, one brief

I ran the audit with a method I wrote for exactly this. One AI model reading a codebase for performance problems finds some of them. Several different models reading the same brief find more, disagree usefully and catch each other’s mistakes. The loop is fixed. Work out which code sits on the slow path. Lock a panel of four or five models from the tools actually installed. Measure the complaint before touching anything. Send the identical read-only brief to every model at once and require findings in a fixed format with file and line evidence. Verify every finding yourself by reading the code behind it. Have the strongest model judge what only the others found and trace the hot paths. Ask the panel to attack the fix plan before any code is written. Fix in batches from lowest risk up, re-measuring after each. Put a different model on every change as a sceptic. Record who caught what.

The panel here was GPT-6 Astra at its highest effort as the lead, Opus 5.5 at high effort, SWE-2 at high effort and Mimo v2.6 Pro at max effort. I was not on the panel. I verified.

FindingAstraOpusSWE-2Mimo
Replay per AI answer grows with session lengthyesyesyesyes
Full copy of the queue sent to the panel per answeryesyesyesyes
Two-second health check redraws the panelyesyesyesyes
Row number inside the drawing keyyesyesyesyes
Honorifics re-stripped on every callyesyesyesrated low
Two-slot AI queue slows burstsyesyesyesyes
Skip the replay when an answer changes nothinground 2
Re-checks triggered by a fold sliding older commentsjudgedyesyes
Folded list re-sent in every updateyes
A display-only switch cancels AI workround 2

Who caught whatTen performance findings, four auditors, one read-only brief. A hollow dot is a finding the auditor reported but rated low; “round 2” came from the lead after it saw the profile

6findings all four auditors made independently

FindingAstraOpusSWE-2Mimo
Replay per AI answer grows with the session P-3
Full copy of the queue sent per answer P-2
Health check redraws the panel P-4
Row number inside the drawing key P-4b
Two-slot AI queue slows bursts P-6
Honorifics re-stripped on every call 52% of CPU
Re-checks after a fold slides comments P-X1judged
Skip the replay when nothing changes P-9round 2
Display-only switch cancels AI work P-X2round 2
Folded list re-sent in every update P-2
Convergence is evidence, and a single-model finding is either the most valuable item on the list or a hallucination. The honorific row is the one the profile had to settle.

Six findings converged across all four, including the honorific re-stripping, which one of them rated low. One more, re-checks after a fold, came from three. Two came from the lead alone, in the second round after it had seen the profile. One came from Mimo alone.

A profile settled the one disagreement

That disagreement is where the baseline measurement paid for itself. A profile, meaning a breakdown of where the processor time went function by function, taken on the rig at 200 comments, put 52% of the time in re-stripping and re-sorting the honorific list on every call, 31% in the word-overlap check that spots near-duplicates, and nothing else above 3%. The finding one auditor had rated low was the single largest cost in the pipeline. The profile settled it, not a vote.

Where the processor time went at 200 commentsShare of processor time by function, from a profile taken on the rig: re-stripping honorifics, and the word-overlap check for near-duplicates

52%of processor time spent re-stripping honorifics, the finding one auditor rated low

Two functions took 83% of the processor time. Both were cacheable, and neither was where the auditors had put the most weight.

Before committing to a plan I built the two caches on their own as a prototype. Interleaved runs on a busy machine gave 12.3 s and 9.6 s for the original against 5.6 s and 4.7 s with the caches, about 2.1 times less processor time. The longest task, which ran between 203 and 366 ms under that load, fell to 46 ms. All 93 matching tests and 46 session tests passed. That was enough to put the caches first.

Seven batches, three that moved the number

The batch plan went to the panel to break before any code was written, and every trap they raised was adopted. Two were about the caches: the honorific list could be changed in place, so it is frozen and the cache is keyed by its contents, and the word cache had no size limit and was shared across tests, so it now drops its oldest entries and clears on reset. One was about sending only changes to the panel: the panel read a missing folded list as an empty one and would have wiped it, so a missing list now means “unchanged”. One was about the shortcut in batch 3: an AI answer of “new question” is not always a no-op, so the shortcut applies only where it provably is, and a bounded sweep keeps follow-up checks flowing.

BatchCPU at 200 / 400 commentsms per comment, last 20Longest taskData to panel
Baseline6,760 / 33,174 ms68.1 / 151.471.4 / 418 ms11.3 / 42.7 MB
0. Feed safety6,893 / 30,634 ms70.0 / 158.772.1 / 184 ms11.3 / 42.7 MB
1. Caches and diffs2,212 / 9,459 ms22.5 / 46.920.3 / 44 ms1.1 / 3.5 MB
2. Panel and idleunchangedunchangedunchangedunchanged
3. Skip needless replays1,014 / 4,750 ms8.9 / 20.820.8 / 47 ms1.1 / 3.5 MB
4. AI queue and safety1,002 / 4,559 ms8.7 / 20.920.2 / 44 ms1.1 / 3.7 MB
5. Models and prompts995 / 4,564 ms8.7 / 20.820.1 / 43 ms1.1 / 3.7 MB
6. Server and packaging1,005 / 4,542 ms8.8 / 20.820.2 / 44 ms1.1 / 3.7 MB
7. Accuracy polish968 / 4,240 ms8.3 / 17.921.8 / 45.5 ms1.2 / 3.7 MB

Batch 0 came first on purpose. It fixed the findings that could put the wrong comment on air or hide a real question, and touched no hot path, so its row looks like the baseline. Correctness before speed.

Batch 1 landed the two caches, dropped a copy-and-sort of records that were already in order, and sent the panel only what changed instead of a full copy of the queue. Processor time fell about three times at both session lengths and the data sent to the panel fell to about a tenth: 42.7 MB to 3.5 MB at 400 comments.

Batch 2 changed only the side panel, so the rig reads it as unchanged. In a real Chromium browser, the number of times the idle sidebar redrew itself over 10 s of health checks fell from 10 to 0.

Batch 3 added the shortcut: when the AI confirms a comment is simply a new question and nothing about it is pending, the extension updates that one card instead of replaying the session. On the rig 64% of answers took it (126 of 197 at 200 comments, 242 of 384 at 400), and calls to the matcher at 200 comments fell from 20,392 to 7,136. Processor time halved again. The simulated AI answers every same-person check “extra question”, which still replays, so the share in production is a measurement to take, not a number to assume.

Batches 4, 5 and 6 changed the AI request queue, the models and the server. None of that is on the processor-time path and the rows say so; the 4 and 5 rows were measured back to back, interleaved with the main branch, in one sitting. Batch 4 did move a different number: with the simulated AI at 3 s, the 10th comment of a burst is now reviewed at 9.25 s instead of 15.25 s.

Batch 7 indexed stored comments by fingerprint, collapsed a three-query check into one and made each panel update look up a row’s on-air button once. The rig sees 3% and 6% less processor time at 200 and 400 comments and 13% less on the last 20 comments. The page-side savings it cannot see.

What stayed the same

The rig, the recorded comments, the 3 s simulated AI and the machine were the same for every row. The test suite and the build were green after every batch. The accuracy check on unseen sessions ran after batches 0, 3 and 5, and the probability check’s count of real questions wrongly hidden stayed at 0 on all four of its roles each time, so the speed work did not buy its numbers by hiding questions.

What this does not show

  • The rig measures processor time outside the browser against a fake page. It cannot see layout, drawing or message-passing costs, so the panel-side findings were checked in Chromium, not with this table.
  • The simulated AI answers in a flat 3 s. Real answers are neither flat nor 3 s, and batch 4’s burst number depends on that assumption.
  • The prototype comparison ran on a busy machine with interleaved runs. The batch table is the idle re-measurement it promised.
  • Everything was measured on an Apple M1. The teacher’s laptop is slower, which is why the 50 ms line mattered at 200 comments rather than 400.
  • Two proposals were recorded as not done: replaying from a saved checkpoint instead of from the start, and splitting a replay into time slices. Both are high risk and only earn their place if the rig still shows tasks over 16 ms after batch 3. It does not, so they stay open.
  • Sending duplicate fallback requests and prompt-cache affinity headers were left out as unverified gains, to be re-evaluated with the bake-off’s latency data.

Evidence

The audit record, the baseline, the profile, the attribution table and the per-batch measurements are in specs/safwa-audit-2026-09-22.md in the Ṣafwa repository, in the sections “Performance record” and “Implementation record”. The repository is private because its test dataset holds viewers’ comments.

This is one of four posts on Ṣafwa. The model comparison that batch 5 waited for is in wrongly hidden questions, 4.0% to 1.9%, the broadcast replay that fixed the confusing cards is in from six confusing cards to none, and the model decision that followed is in from 41 hidden questions to zero.