Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

CTC and RNN-T

Frame-synchronous models with an alignment you collapse

A recogniser reads a sequence of audio frames and must produce a shorter sequence of letters, and nobody tells it which frames belong to which letter. CTC handles that by adding a blank symbol to the vocabulary. The network produces one distribution per frame; a blank means "nothing is being said here", and the training objective sums over every monotonic path that collapses to the correct transcript, so the alignment is marginalised out rather than supervised. At inference you take the argmax per frame, merge runs of the same symbol, and drop the blanks. The heatmap below is exactly that posterior picture: a bright spike wherever a symbol is being spoken, a blank row for the gaps.

The catch is that CTC assumes the frames are conditionally independent given the transcript, so it has no internal language model and no memory of what it just emitted. RNN-T fixes that with a predictor — a small language model over the tokens emitted so far — whose state is combined with each encoder frame by a joint network. At every cell $(t, u)$ the joint network predicts either a blank, which advances the time index and emits nothing, or the next token, which advances the predictor instead. That is a transducer, and it is naturally streaming: tokens can be emitted as audio arrives, without waiting for the whole utterance and without CTC's independence assumption.

Per-frame posteriors drawn with Guide.drawHeatmap. The top row is the blank symbol; the bright diagonal band is the alignment CTC discovers on its own.

The RNN-T lattice. Moving right is a blank (advance the encoder); moving up emits a token (advance the predictor). Every path is monotonic, so decoding can begin before the sentence ends.

💡 By the end of this part you'll be able to explain what the CTC blank buys and how the collapse rule works, describe what the RNN-T predictor adds, say why a Whisper-style encoder cannot stream and what a block-causal mask changes, and account for every millisecond of a full-duplex turn budget.
2

Whisper and the streaming problem

Strong offline, awkward live

Whisper replaced the frame classifier with an attention encoder-decoder trained on hundreds of thousands of hours of weakly labelled audio. Audio is resampled to 16 kHz, turned into a log-mel spectrogram, and then padded or cropped to a fixed thirty-second window. The encoder attends over all fifteen hundred frames at once, in both directions, and the decoder cross-attends to that representation while writing text. The fixed window and the bidirectional attention are what make the model robust and easy to train on a messy corpus; they are also exactly what make it non-streaming. A frame near the start of the window already knows about a frame near the end.

To run it online you have to decide how much context is available, and both extremes are bad. Wait for the full thirty seconds and the latency is thirty seconds. Cut the audio into chunks and decode each independently and the model loses context and mangles any word that straddles a boundary, which is a well-documented word-error-rate cliff. The middle ground is a mask. A block-causal mask lets each frame attend to every frame in its own block and all earlier blocks, but to nothing later, so a fixed look-ahead of one or two blocks buys back most of the accuracy at a bounded delay. The heatmap below is the two masks side by side: full bidirectional, where every frame sees every other, and block-causal, where the staircase of zeros is the future the model is not allowed to see.

The fixed thirty-second window on top, and the encoder attention mask below. Rows are frames being encoded, columns are frames they attend to; the dark triangle in block-causal mode is the unavailable future.

⚠ Re-encoding is the hidden cost. A naive chunked Whisper re-runs the encoder on a fresh thirty-second window for every chunk, so the compute per second of audio is many times the offline figure. Caching the encoder state, or training with a block-causal mask in the first place, is what makes streaming affordable rather than merely possible.
3

The two-pass fix

Cheap partials, then a rescoring pass

The standard answer is to stop pretending one pass can be both fast and accurate. U2 and U2++ train a single hybrid model with two decoders: a first pass that is causal or chunk-based and emits partial hypotheses with low latency, and a second pass whose attention decoder re-decodes and rescores those hypotheses using more right context. The streaming latency is a knob — the chunk size — and the final output approaches the non-streaming accuracy because the second pass gets to look ahead. The same idea appears in a transformer encoder as the block-causal mask from the previous step: attention restricted to the current block plus preceding blocks, so the model runs online with a fixed look-ahead instead of an unbounded one.

Note what is and is not being bought. A two-pass model does not remove the latency; it splits it into a small partial latency, which is what the interface shows while the user is still speaking, and a larger finalisation latency, by which time the text is as good as offline. Interleaving full-attention and causal layers is a third variant of the same trade: a few bidirectional layers recover some global context while the rest of the stack stays causal. In every version, context is the currency and latency is the price.

Chunks of audio, the causal first pass that emits partials one chunk behind the speech, and the second pass that re-decodes each chunk once its look-ahead has arrived.

💡 Latency is a control, not a property. The same weights serve a low-latency partial and a near-offline final transcript; the chunk size and the look-ahead decide which you get. That is why the two-pass design became the default in production recognisers.
4

The realtime duplex budget

Voice activity, barge-in, and the reply the user hears

A voice agent layers turn-taking on top of recognition. A voice-activity detector separates speech from silence, endpointing decides when a turn is over, and barge-in is the case where the user starts talking while the assistant is still speaking — which requires the agent to stop, listen and revise. In a cascade — voice activity, then recognition, then the language model, then synthesis — every stage is a separate service and the latency is the sum. The number the user actually feels is the gap between the end of their speech and the first sound of the reply, and in a well-built cascade the recogniser runs concurrently with the speech, so what remains in the gap is the network hop, the time to first token and the TTS first chunk.

The alternative is a native speech-to-speech model. Moshi models the user's audio and its own audio as two parallel streams inside one transformer and removes explicit speaker turns altogether, reporting a theoretical latency of about 160 ms and roughly 200 ms in practice. Qwen2.5-Omni keeps a thinker that writes text and a talker that emits speech tokens from the thinker's hidden states, and stream-decodes with a sliding-window decoder to cut the first-package delay. Both fold the text bottleneck away, at the cost of intermediate text you can no longer inspect and control. The simulator below makes the cascade arithmetic concrete: push the stages around, then flip the barge-in switch and watch a pending reply get cancelled.

A turn from GenMedia.timeline.turn: user speech, streaming recognition overlapping its end, the reply gap in its parts, and the assistant. The barge-in marker cancels a reply that has not started.

The full cascade chain assembled with GenMedia.timeline.budget and drawn by Guide.drawStacked. The dashed line is the 300 ms figure a natural turn is allowed.

A native duplex model does not remove these stages; it runs them jointly, so the recogniser's finalise time stops sitting in front of the reply.

⚠ Two hundred to three hundred milliseconds is the whole budget. Below about 200 ms a reply feels like a reflex; past roughly 300 ms listeners read it as hesitation, and past a second they start talking again. That is why barge-in handling is not a nice-to-have: an agent that cannot be interrupted will be interrupted anyway, badly.
5

Where this shows up

Recognition is the front half of every voice agent

Serving

Realtime voice as a workload

A duplex conversation is a serving problem before it is a modelling one: streaming requests, tiny batches, a hard latency objective and a prefix cache that turns over every turn. The workload-page treatment of realtime voice is in LLM Serving, Part 18, and the numbers there are the same ones this page's budget bar is spending.

Sound · Synthesis

The answering half

Everything the reply spends after the first token is text-to-speech: the codec vocabulary, the autoregressive first chunk and the vocoder. Recognition and synthesis share the token vocabulary from the neural codec, which is why a native duplex model can run both in one sequence.

The frontier question is no longer whether recognition can be accurate, but whether the two halves of a conversation can be one model. Cascade systems are measurable, debuggable and easy to ship; native duplex models are faster and handle overlap, interruption and emotion that a turn-based pipeline simply throws away. Both are being deployed, and the part that follows is about the problem that arrives as soon as either one is good enough to fool a listener: telling generated speech from recorded speech, and attributing it.

Further reading

The recognition literature is unusually well marked by a handful of papers. CTC makes the alignment marginalisable, RNN-T makes it a streaming transducer, Whisper makes attention the default and streaming the problem, U2 and U2++ supply the two-pass answer, and Moshi and Qwen2.5-Omni are the current argument that the whole pipeline should be one model.

Graves and coauthors introduce the blank symbol and the collapse rule; Graves follows with the transducer; Radford and coauthors scale the encoder-decoder; Zhang and Wu and coauthors unify streaming and non-streaming with a rescoring second pass; and Défossez and Xu and coauthors remove the turns entirely.

Cheat sheet

TermMeaning here
CTCA frame classifier trained without alignment; the blank symbol absorbs the slack
Collapse ruleMerge runs of the same symbol, then drop blanks; a repeat needs a blank between
RNN-TA transducer: a predictor over emitted tokens joined to each encoder frame
Blank (transducer)No new token at this frame; advance time instead of the predictor
Log-melThe 16 kHz spectrogram Whisper's encoder consumes
Thirty-second windowWhisper pads or crops every chunk to it; the irreducible offline latency
Bidirectional attentionEvery frame sees every other frame; the reason Whisper is not causal
Block-causal maskAttention to the current block and all earlier ones; a bounded look-ahead
Two-pass / U2Causal first pass emits partials; a second pass rescores with right context
Chunk sizeThe streaming latency knob; partials land one chunk behind the speech
VADVoice activity detection; separates speech from silence
EndpointingDeciding the turn is over, which is where most of the gap is spent
Barge-inThe user talks over the assistant; the agent must stop and listen
CascadeVAD, ASR, LLM and TTS as separate stages, summed at the user
Native duplexOne model over both audio streams; Moshi and Qwen2.5-Omni
Felt gapSpeech end to first reply sound: network + TTFT + TTS first chunk
7

Check your understanding

0/4 answered