Batching a token stream to one repaint per frame

A streamed answer arrives as hundreds of deltas. Rendering each one is hundreds of React commits for text nobody has read yet. Collapsing them to one paint per animation frame is four lines, and it moves the hard part somewhere else.

Many thin vertical ticks arriving in a dense row, collapsing into a small number of evenly spaced frames

A model streaming a reply emits deltas far faster than a screen can show them. Render on every delta and you get one React commit per token, most of them painting text that is replaced a few milliseconds later. The fix is four lines. The reason to write about it is what the four lines push into the rest of the code.

The problem

The chat surface consumes a server-sent event stream carrying several kinds of event: reasoning text, answer tokens, and lifecycle events for each tool the agent runs. They are interleaved, because the agent thinks, calls a tool, thinks again.

The naive consumer calls a state setter per event. That is correct and it is wasteful. Each setter schedules a render, each render reconciles a growing timeline, and the work scales with the length of the answer rather than with anything the reader perceives. Nobody can read at the rate the tokens arrive.

The usual reach is a debounce or a fixed-interval timer. Both are worse than they look. A debounce delays the last chunk of an answer until the stream goes quiet, which is exactly when the reader is waiting for it. A timer at 60ms is either slower than the display or faster, and it is never aligned to when the browser actually paints.

What batching to a frame does

The display has a natural clock, and the browser will tell you when it is:

let streamRafId: number | null = null;

const flushTimeline = () => {
  streamRafId = null;
  syncLive();
};

const scheduleTimelineFlush = () => {
  if (streamRafId !== null) return;
  streamRafId = requestAnimationFrame(flushTimeline);
};

Every event mutates a plain array and calls scheduleTimelineFlush. The guard means the second through hundredth call in a frame do nothing at all. React renders once per frame, with everything that arrived in it.

ApproachRenders per secondLatency of the last tokenAligned to paint
Setter per eventAs fast as the streamNoneNo
Debounce 60msFewerUp to 60ms, worst at the endNo
Fixed intervalFixedUp to the intervalNo
One per animation frameAt most the refresh rateUnder one frameYes

The last row is not a compromise between the others. Painting more often than the display refreshes is work with no observable result, so the frame rate is the correct ceiling rather than a tuned number.

This generalises past chat. Any high-frequency source rendered for a human has this shape: a cursor position stream, log tailing, a progress feed. The rule is that the render rate should be set by the display, and the data rate should only decide how much each render has to catch up on.

What it moves, rather than solves

Batching means the array is mutated many times between renders, so the state has to be correct after arbitrary interleaving rather than after each event. That work lives in a 31-line module of pure functions, deliberately separated from the component:

export function appendTimelineText(timeline: TimelineEntry[], text: string): void {
  const last = timeline[timeline.length - 1];
  if (last && last.kind === 'text') last.text += text;
  else timeline.push({ kind: 'text', text });
}

Streamed prose merges into the trailing text node, unless a tool step was pushed since the last one, in which case it starts a new node. That is what keeps reasoning and tool steps interleaved in arrival order instead of collecting into two separate blocks.

The step lifecycle is the subtler one. A tool can be called more than once in a turn, so closing a step cannot match by tool name alone:

const step = steps.find(s => s.tool === tool && !s.done && !s.error);
if (!step) return;

It closes the oldest still-running step for that tool. Two geocodes in one turn each get their own spinner and each closes in order.

Both functions take an array and return nothing, which means they are testable without a stream, a server, or a component. That is most of the value of pulling them out.

Where it breaks

The early return is silent. if (!step) return; handles a close event naming a tool with no open step, which is the right call in a UI: never crash the interface over a stray event. It also means a genuine protocol bug produces no symptom at that layer. This bit during a refactor that moved tool execution below the layer emitting the close events. Four spinners opened, none closed, and the errors were reported against a tool with no spinner, so they were dropped too. The stream was wrong and the UI silently absorbed it. A development-mode warning on an unmatched close would have named it immediately.

Mutating an array behind a setter is a deliberate rule break. syncLive calls setLiveTimeline([...timeline]), so a fresh array reaches React each flush, but the entries inside are the same objects, mutated in place. Nothing enforces that no component holds a stale reference to one. It is fast and it works; it is not safe by construction.

One flush per frame is a ceiling, not a floor. If a single flush becomes expensive enough to overrun a frame, the scheduling gives you no protection: the next frame's work simply queues. Batching bounds how often you render, never how long a render takes.

Nothing here is cancelled on unmount. If a turn is still streaming when the view goes away, a scheduled frame can call a setter on a component that no longer exists. In practice the stream is torn down first, but that ordering is incidental rather than enforced.

More on what the system does with a question and how to read the answer it streams back. FRED itself is at biofred.us if you would rather watch it stream than read about it.

References

  1. MDN. Window.requestAnimationFrame().

  2. MDN. Server-sent events.

  3. React. Queueing a series of state updates.