Speech to text, and answering out loud
Recognition is synthesis run backwards, and it inherits the same alignment problem with one extra constraint: it has to run while the user is still talking. The field solved that in two jumps — CTC showed that a frame classifier could learn an alignment it was never given, and RNN-T turned recognition into a streaming transducer. Then Whisper raised accuracy with a fixed thirty-second window attended in both directions, which is precisely the design that refuses to stream. This part follows that arc and lands on the thing a user actually feels: the few hundred milliseconds between the end of their sentence and the first sound of the reply.
CTC and RNN-T
Frame-synchronous models with an alignment you collapse
A recogniser reads a sequence of audio frames and must produce a shorter sequence of letters, and nobody tells it which frames belong to which letter. CTC handles that by adding a blank symbol to the vocabulary. The network produces one distribution per frame; a blank means "nothing is being said here", and the training objective sums over every monotonic path that collapses to the correct transcript, so the alignment is marginalised out rather than supervised. At inference you take the argmax per frame, merge runs of the same symbol, and drop the blanks. The heatmap below is exactly that posterior picture: a bright spike wherever a symbol is being spoken, a blank row for the gaps.
The catch is that CTC assumes the frames are conditionally independent given the transcript, so it has no internal language model and no memory of what it just emitted. RNN-T fixes that with a predictor — a small language model over the tokens emitted so far — whose state is combined with each encoder frame by a joint network. At every cell $(t, u)$ the joint network predicts either a blank, which advances the time index and emits nothing, or the next token, which advances the predictor instead. That is a transducer, and it is naturally streaming: tokens can be emitted as audio arrives, without waiting for the whole utterance and without CTC's independence assumption.
Per-frame posteriors drawn with Guide.drawHeatmap. The top row is the blank symbol; the bright diagonal band is the alignment CTC discovers on its own.
The RNN-T lattice. Moving right is a blank (advance the encoder); moving up emits a token (advance the predictor). Every path is monotonic, so decoding can begin before the sentence ends.
Whisper and the streaming problem
Strong offline, awkward live
Whisper replaced the frame classifier with an attention encoder-decoder trained on hundreds of thousands of hours of weakly labelled audio. Audio is resampled to 16 kHz, turned into a log-mel spectrogram, and then padded or cropped to a fixed thirty-second window. The encoder attends over all fifteen hundred frames at once, in both directions, and the decoder cross-attends to that representation while writing text. The fixed window and the bidirectional attention are what make the model robust and easy to train on a messy corpus; they are also exactly what make it non-streaming. A frame near the start of the window already knows about a frame near the end.
To run it online you have to decide how much context is available, and both extremes are bad. Wait for the full thirty seconds and the latency is thirty seconds. Cut the audio into chunks and decode each independently and the model loses context and mangles any word that straddles a boundary, which is a well-documented word-error-rate cliff. The middle ground is a mask. A block-causal mask lets each frame attend to every frame in its own block and all earlier blocks, but to nothing later, so a fixed look-ahead of one or two blocks buys back most of the accuracy at a bounded delay. The heatmap below is the two masks side by side: full bidirectional, where every frame sees every other, and block-causal, where the staircase of zeros is the future the model is not allowed to see.
The fixed thirty-second window on top, and the encoder attention mask below. Rows are frames being encoded, columns are frames they attend to; the dark triangle in block-causal mode is the unavailable future.
The two-pass fix
Cheap partials, then a rescoring pass
The standard answer is to stop pretending one pass can be both fast and accurate. U2 and U2++ train a single hybrid model with two decoders: a first pass that is causal or chunk-based and emits partial hypotheses with low latency, and a second pass whose attention decoder re-decodes and rescores those hypotheses using more right context. The streaming latency is a knob — the chunk size — and the final output approaches the non-streaming accuracy because the second pass gets to look ahead. The same idea appears in a transformer encoder as the block-causal mask from the previous step: attention restricted to the current block plus preceding blocks, so the model runs online with a fixed look-ahead instead of an unbounded one.
Note what is and is not being bought. A two-pass model does not remove the latency; it splits it into a small partial latency, which is what the interface shows while the user is still speaking, and a larger finalisation latency, by which time the text is as good as offline. Interleaving full-attention and causal layers is a third variant of the same trade: a few bidirectional layers recover some global context while the rest of the stack stays causal. In every version, context is the currency and latency is the price.
Chunks of audio, the causal first pass that emits partials one chunk behind the speech, and the second pass that re-decodes each chunk once its look-ahead has arrived.
The realtime duplex budget
Voice activity, barge-in, and the reply the user hears
A voice agent layers turn-taking on top of recognition. A voice-activity detector separates speech from silence, endpointing decides when a turn is over, and barge-in is the case where the user starts talking while the assistant is still speaking — which requires the agent to stop, listen and revise. In a cascade — voice activity, then recognition, then the language model, then synthesis — every stage is a separate service and the latency is the sum. The number the user actually feels is the gap between the end of their speech and the first sound of the reply, and in a well-built cascade the recogniser runs concurrently with the speech, so what remains in the gap is the network hop, the time to first token and the TTS first chunk.
The alternative is a native speech-to-speech model. Moshi models the user's audio and its own audio as two parallel streams inside one transformer and removes explicit speaker turns altogether, reporting a theoretical latency of about 160 ms and roughly 200 ms in practice. Qwen2.5-Omni keeps a thinker that writes text and a talker that emits speech tokens from the thinker's hidden states, and stream-decodes with a sliding-window decoder to cut the first-package delay. Both fold the text bottleneck away, at the cost of intermediate text you can no longer inspect and control. The simulator below makes the cascade arithmetic concrete: push the stages around, then flip the barge-in switch and watch a pending reply get cancelled.
A turn from GenMedia.timeline.turn: user speech, streaming recognition overlapping its end, the reply gap in its parts, and the assistant. The barge-in marker cancels a reply that has not started.
The full cascade chain assembled with GenMedia.timeline.budget and drawn by Guide.drawStacked. The dashed line is the 300 ms figure a natural turn is allowed.
A native duplex model does not remove these stages; it runs them jointly, so the recogniser's finalise time stops sitting in front of the reply.
Where this shows up
Recognition is the front half of every voice agent
Realtime voice as a workload
A duplex conversation is a serving problem before it is a modelling one: streaming requests, tiny batches, a hard latency objective and a prefix cache that turns over every turn. The workload-page treatment of realtime voice is in LLM Serving, Part 18, and the numbers there are the same ones this page's budget bar is spending.
The answering half
Everything the reply spends after the first token is text-to-speech: the codec vocabulary, the autoregressive first chunk and the vocoder. Recognition and synthesis share the token vocabulary from the neural codec, which is why a native duplex model can run both in one sequence.
The frontier question is no longer whether recognition can be accurate, but whether the two halves of a conversation can be one model. Cascade systems are measurable, debuggable and easy to ship; native duplex models are faster and handle overlap, interruption and emotion that a turn-based pipeline simply throws away. Both are being deployed, and the part that follows is about the problem that arrives as soon as either one is good enough to fool a listener: telling generated speech from recorded speech, and attributing it.
Further reading
The recognition literature is unusually well marked by a handful of papers. CTC makes the alignment marginalisable, RNN-T makes it a streaming transducer, Whisper makes attention the default and streaming the problem, U2 and U2++ supply the two-pass answer, and Moshi and Qwen2.5-Omni are the current argument that the whole pipeline should be one model.
Graves and coauthors introduce the blank symbol and the collapse rule; Graves follows with the transducer; Radford and coauthors scale the encoder-decoder; Zhang and Wu and coauthors unify streaming and non-streaming with a rescoring second pass; and Défossez and Xu and coauthors remove the turns entirely.
- Alex Graves, Santiago Fernández, Faustino Gomez and Jürgen Schmidhuber, "Connectionist Temporal Classification: Labelling Unsegmented Sequence Data with Recurrent Neural Networks", 2006 — the blank symbol and the alignment marginalised out.
- Alex Graves, "Sequence Transduction with Recurrent Neural Networks", 2012 — the RNN-T, with a predictor feeding a joint network.
- Alec Radford, Jong Wook Kim, Tao Xu and coauthors, "Robust Speech Recognition via Large-Scale Weak Supervision", 2022 — Whisper, the fixed thirty-second window and the encoder-decoder.
- Binbin Zhang, Di Wu, Zhuoyuan Yao and coauthors, "Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition", 2020 — U2, hybrid CTC and attention with a rescoring second pass.
- Di Wu, Binbin Zhang, Chao Yang and coauthors, "U2++: Unified Two-pass Bidirectional End-to-end Model for Speech Recognition", 2021 — the bidirectional two-pass refinement.
- Alexandre Défossez, Laurent Mazaré, Manu Orsini and coauthors, "Moshi: a speech-text foundation model for real-time dialogue", 2024 — full-duplex speech-to-speech with parallel user and model streams.
- Jin Xu, Zhifang Guo, Jinzheng He and coauthors, "Qwen2.5-Omni Technical Report", 2025 — a thinker writing text and a talker emitting speech, stream-decoded.
Cheat sheet
| Term | Meaning here |
|---|---|
| CTC | A frame classifier trained without alignment; the blank symbol absorbs the slack |
| Collapse rule | Merge runs of the same symbol, then drop blanks; a repeat needs a blank between |
| RNN-T | A transducer: a predictor over emitted tokens joined to each encoder frame |
| Blank (transducer) | No new token at this frame; advance time instead of the predictor |
| Log-mel | The 16 kHz spectrogram Whisper's encoder consumes |
| Thirty-second window | Whisper pads or crops every chunk to it; the irreducible offline latency |
| Bidirectional attention | Every frame sees every other frame; the reason Whisper is not causal |
| Block-causal mask | Attention to the current block and all earlier ones; a bounded look-ahead |
| Two-pass / U2 | Causal first pass emits partials; a second pass rescores with right context |
| Chunk size | The streaming latency knob; partials land one chunk behind the speech |
| VAD | Voice activity detection; separates speech from silence |
| Endpointing | Deciding the turn is over, which is where most of the gap is spent |
| Barge-in | The user talks over the assistant; the agent must stop and listen |
| Cascade | VAD, ASR, LLM and TTS as separate stages, summed at the user |
| Native duplex | One model over both audio streams; Moshi and Qwen2.5-Omni |
| Felt gap | Speech end to first reply sound: network + TTFT + TTS first chunk |