Omni models and the real-time loop
The last step is to put every modality into one model and run it live. A modality here means one kind of data — written text, an image, a recording of speech — and an omni model is a single model that handles all of them together instead of handing each one to its own specialist sub-model. That means text tokens and audio tokens in a single interleaved stream, a running sequence in which the two kinds of token alternate rather than sitting in separate pipelines: picture two people passing notes in a relay, each one handing the next a short message in the order it arrives, so no single person ever holds the whole conversation. It also means an input stream and an output stream open at the same time, so the model is still hearing the user while it is already speaking, and a reply that begins while the user is still able to interrupt it. None of that is a modelling flourish: full duplex — both directions live at once, the way a phone call works but a walkie-talkie does not — turns the whole problem into a latency budget, a running total of the milliseconds spent before the reply can be heard. Every stage — recognition, network, the first text token, the first audio token — spends from that budget, and the budget is only the few hundred milliseconds a listener will tolerate before the pause stops feeling like a conversation. This part assembles the loop and does the accounting.
One model, many streams
Interleaved tokens in one sequence
An omni model does not concatenate modalities at the edges — gluing a transcript on to a description, say — it interleaves them, so a single sequence alternates text subwords, image patches and codec audio tokens in whatever order the task requires. A subword is one word piece from a text tokeniser; a patch is one small tile an image is cut into; and a codec audio token is one symbol from a fixed list that a neural codec — a pair of networks that compresses sound down to a compact code and reconstructs it again — emits for a short slice of audio. Interleaving means all three kinds of token live in one running sequence, one after another, in the order they actually occurred. The training objective is unchanged — predict the next token — but the vocabulary is shared, meaning the same fixed list of symbols covers every modality, and the model has to learn when a text token should follow an audio token and when it should not. That single sequence is what lets a model answer a question about a recording, quote a phrase it just heard, or switch from describing a picture to reading a caption aloud, without a routing layer deciding which sub-model to call.
The cost is a shared budget again, this time for output as well as input: audio tokens are far more numerous per second than text, so a stream that talks spends its context on sound. Context here means the fixed window of tokens the model can hold in mind at once, and once that window is full something has to be dropped, so every audio token the reply spends is a token the reasoning cannot use. Concretely, one second of speech is hundreds of codec tokens, while a sentence of the same duration is a few dozen subwords, which is why a spoken reply costs far more of the window than the same answer typed out. The strip below shows the interleaving as a token sequence, with the slider setting how many audio tokens accompany each text token.
One interleaved token stream, wrapped onto two rows. Each square is a token; the slider sets how many audio tokens accompany each text token, which is what makes speech-heavy output so expensive.
The thinker and the talker
Text for reasoning, audio for delivery
One way to keep the quality of a language model while producing speech natively is to split the decoding into two stages that share the same trunk. The trunk is the shared backbone network — the stack of transformer layers that both stages read from — and decoding is the act of generating output one token at a time. A thinker decodes text tokens as an ordinary language model, which is where the reasoning, the factual grounding and the instruction following live; grounding here means tying the words to facts and to the input rather than letting the language model drift on its own. A talker takes the thinker's hidden states — the internal vectors the thinker computes at each position, not just the tokens it finally samples out — and emits audio codec tokens for the reply. Those hidden states are richer than the finished text, because they carry the thinker's partial intent before it has settled into words. The talker is small relative to the thinker, and because it conditions on hidden states — that is, uses them as its input — rather than on finished text, it can start before the sentence is complete.
The split buys three things. First, text stays inspectable and controllable, since you can read what the model "meant" to say and catch a bad plan before it is spoken. Second, the audio rate can be chosen independently of the text rate, so the talker can stream — emit its tokens incrementally as they are produced, instead of waiting for the whole reply and releasing it in one block — in whatever codec the deployment uses. Third, and most important for a conversation, the two can overlap: the thinker writes ahead while the talker catches up, which shortens the time to the first audible syllable — the number the user actually feels. A system that waited for the complete sentence before starting to speak would pay for every extra word in the thinker's plan, and that added delay is exactly what makes a reply feel sluggish.
The thinker-and-talker split: audio tokens in, a text decoder that reasons, and a small audio decoder reading the thinker's hidden states to emit speech tokens in parallel.
A single-model alternative exists: one transformer emitting text and audio tokens in one sequence, as Moshi does with parallel user and model streams. The split is a practical compromise that keeps a strong text backbone intact.
Speaking while listening
Full duplex and the interrupt
A turn-based system alternates: listen, decode, answer. This strict turn-taking is the walkie-talkie pattern — one side talks, then the other, and nobody may speak until the channel is handed over. A full-duplex model keeps both streams open, so it can be listening for a new utterance (one stretch of speech from the user) while it is still producing the current reply. That is what makes barge-in work, barge-in being the user cutting in before the model has finished. When the user starts talking mid-reply, the model must not merely stop — it must keep the audio it heard, revise its plan and respond to the interruption. This is harder than it sounds, because the microphone is also picking up the model's own voice: the system has to separate the user's speech from its own output, which is the job usually called echo cancellation, and it has to notice when a human voice has actually started, which is what voice activity detection does. A turn-based pipeline handles an interrupt poorly, because the interrupt arrives while its synthesis stage — the stage that turns the planned answer into sound — is mid-sentence and its recogniser, the model that turns incoming audio into words, has not been running.
The timeline below models one turn: the user's speech, recognition overlapping its end, the reply gap assembled from the network hop, the time to first text token and the time to first audio token, and the assistant's reply. The reply gap is the stretch a listener actually hears as silence, and it is built from stages that each cost time: the network hop is the round trip that carries the audio to the server, the time to the first text token is how long the thinker takes to produce its first word, and the time to the first audio token is how long the talker then needs to turn that word into sound. Flip the barge-in switch and a reply that has not started yet is cancelled, which is the cheap case; a reply already in flight — audio already playing — is the expensive one, because the model has committed to a sentence it now has to abandon.
One turn from GenMedia.timeline.turn, with the reply gap's stages assembled by GenMedia.timeline.budget. The barge-in marker cancels a reply that has not yet begun.
The realtime budget
Everything on one GPU
The latency a listener feels is not the sum of the pipeline's work; it is the gap between the end of their speech and the first sound of the reply. That felt gap is what "real-time" means here: not that the whole answer is instantaneous, but that the reply begins soon enough to read as an answer rather than a delay. Streaming recognition overlaps the speech — it emits words while the user is still talking — so it does not sit in that gap, but the network hop, the time to the first text token and the time to the first audio token all do, and together they have to fit inside roughly 200–300 ms. Why that number? Below it the exchange reads as a reflex; above it the pause starts to read as hesitation, or as a connection that has dropped. Under load that arithmetic breaks, because the audio decoders, the recogniser and the thinker share one accelerator — the single GPU doing all of the computation — and every extra concurrent conversation, another user being served at the same moment, multiplies the work on the same silicon.
The budget bar below sums the stages and scales the compute-bound ones — the stages whose cost grows with how much work the GPU is being asked to do — by the load slider, so you can watch a comfortable budget cross the 300 ms line without any single stage becoming unreasonable. That crossing is the reason serving a real-time voice model is a scheduling problem as much as a modelling one: the model can be perfectly accurate and still feel broken if the server cannot admit the next conversation without pushing this one past the limit. Scheduling here means deciding which requests share the accelerator and when, and that decision shows up directly in the pause the user hears.
The reply budget from GenMedia.timeline.budget, drawn with Guide.drawStacked. The dashed line is the 300 ms target; the load slider scales every compute-bound stage together.
Where this shows up
The loop, closed
One model, no cascade
Moshi models the user's audio and its own audio as parallel streams in one transformer and drops explicit speaker turns — that is, it never waits for a formal hand-off between the two sides, so both voices sit in the same sequence at once. Qwen2.5-Omni pairs a thinker writing text with a talker emitting speech tokens. Both are the architecture this part has been assembling, and both are described in more detail alongside the recognition machinery in the audio volume.
Realtime voice as a workload
A duplex conversation is a serving problem before it is a modelling one: long-lived streaming requests, very small batches, a hard tail-latency objective and a prefix cache that turns over every turn. A prefix cache is the model's running notebook: it keeps a note of the tokens it has already processed so it does not redo that work when the next token arrives, and because a conversation keeps adding turns, that notebook is rewritten constantly rather than held steady. Serving voice means holding many such notebooks open at once, each one small in batch size but long in duration. The workload treatment of realtime voice, with the same budget numbers, is LLM Serving, workloads.
That closes the arc. It began with cutting an image into patches and ends with one trunk reading text, images, video and audio in a single interleaved sequence while a conversation runs at the speed of speech. Each earlier idea was a prerequisite for this one: without a shared embedding space the modalities could not be compared at all, without the bridge there was no first way to connect them, and without knowing the token cost of an image, a clip or a second of audio there would be no way to fit the streams into one budget. The pieces in between — the shared space, the bridge, resolution, hallucination, video's token arithmetic, audio's two front ends — are all in service of this loop.
Further reading
The omni literature is young and moves fast, which is why the list below is short and recent rather than canonical. These papers cover the two-stream speech-text model — one transformer holding the user's voice and the model's voice at once — the thinker-talker split, and the first attempts to put speech into a language model's own token space, so that listening and speaking use the same kind of symbol as reading and writing.
- Alexandre Défossez, Laurent Mazaré, Manu Orsini and coauthors, "Moshi: a speech-text foundation model for real-time dialogue", 2024 — full-duplex speech-to-speech with parallel user and model streams in one transformer.
- Jin Xu, Zhifang Guo, Jinzheng He and coauthors, "Qwen2.5-Omni Technical Report", 2025 — a thinker writing text and a talker emitting speech tokens, stream-decoded for low first-package latency.
- Dong Zhang, Shimin Li, Xin Zhang and coauthors, "SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities", 2023 — discrete speech units placed in a language model's token space.
- Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen and coauthors, "AudioPaLM: A Large Language Model That Can Speak and Listen", 2023 — one model over joint text and audio token vocabularies.
- Qingkai Fang, Shoutao Guo, Yan Zhou and coauthors, "LLaMA-Omni: Seamless Speech Interaction with Large Language Models", 2024 — the thinker-talker decomposition for low-latency spoken replies.
Cheat sheet
| Term | Meaning here |
|---|---|
| Interleaved stream | One sequence mixing text, image and audio tokens under a next-token objective |
| Thinker | The text decoder where reasoning, grounding and instruction following live |
| Talker | A smaller audio decoder that reads the thinker's hidden states and emits codec tokens |
| Full duplex | Input and output streams open at once, so the model can listen while it speaks |
| Barge-in | The user interrupts; the model stops, keeps what it heard and revises |
| Felt latency | The gap from the end of the user's speech to the first sound of the reply |
| TTFT vs TTFA | Time to the first text token versus the first audio token; the talker's extra hop |
| The budget | Roughly 200–300 ms, set against a tail percentile, not an average |