Text to speech
Text is a sequence of symbols with no duration; speech is a waveform with a great deal of duration. Nearly everything in a speech synthesiser is a way of deciding how long each symbol lasts and what it sounds like, and the modern answer has moved from a hand-aligned spectrogram to a language model over codec tokens. This part follows that arc: attention alignment and the monotonicity it had to be forced into, the explicit duration models that replaced it, the codec language model that made zero-shot cloning routine, and the two-stage autoregressive-plus-diffusion design that trades a little complexity for streaming latency.
From text to sound
Five stages, and one alignment in the middle
The pipeline below is the shape almost every modern system shares. Text is normalised and turned into symbols. An encoder gives each symbol a vector. An alignment stage expands those vectors along the time axis, deciding how many frames of sound each symbol owns. An acoustic model produces a representation of sound — historically a mel spectrogram, now usually codec tokens — and a vocoder or decoder turns that representation into a waveform. The alignment stage is the one that has no counterpart in machine translation, and it is where the design battles of the last decade were fought.
The reason alignment is hard is that it is a many-to-many correspondence with no labels at training time. We know the text and we know the audio, and we do not know which frames belong to which phoneme. Early systems learned it implicitly with attention, which discovers a soft monotonic map as a side effect of predicting spectrogram frames. Later systems predicted it explicitly with a duration model, which is faster and more controllable. The codec language models of the last few years sidestep explicit alignment almost entirely by treating speech as a token sequence and letting a strong language model learn the correspondence from data.
The text-to-speech pipeline. The middle box — alignment and duration — is the stage the rest of this part is about.
Read the diagram left to right. Text becomes vectors, vectors are stretched in time, the stretched sequence becomes sound tokens, and the tokens become a waveform.
Alignment and duration
Which frames belong to which symbol
Attention-based synthesisers such as Tacotron let the decoder attend over the encoder's symbol vectors while generating frames one at a time. Trained on enough data, the attention settles into a roughly diagonal band: each frame looks at the symbol currently being spoken, and the band moves left to right as the utterance progresses. That band is the alignment, learned without labels. It is also fragile. Because the map is soft, the decoder can skip a symbol, repeat it, or stall, and failures are catastrophic rather than gradual.
Explicit duration models replaced the soft map with a hard one. A duration predictor estimates, for each symbol, how many frames it spans; at inference those counts are used to expand the symbol sequence to frame length, and the acoustic model becomes a non-autoregressive mapping from the enlarged sequence to spectrograms. This is the design of FastSpeech and its descendants. It is fast, it parallelises, and it makes timing directly controllable — you can slow a word down by editing its duration. The cost is that a duration error becomes an audible timing error rather than something attention can smooth away.
The heatmap is a hard alignment of the kind a duration model produces: a bright diagonal band, one frame column per time step, one row per symbol. Move the sliders to change how many symbols there are and how long the utterance runs; the band steepens or flattens but stays monotonic, because speech does not go backwards.
A monotonic alignment matrix: rows are text symbols, columns are frames. Bright means the symbol is being spoken in that frame.
Codec language models
Speech as text
Once a codec has reduced speech to a token stream, synthesis becomes conditional language modelling: given the text tokens, predict the acoustic tokens. VALL-E framed it exactly this way and added the property that made the field take notice — zero-shot cloning. Because the model is conditioned on both the text and a short acoustic prompt, three seconds of a target speaker's audio are enough to make it continue in that voice, with no fine-tuning and no per-speaker parameters. The prompt supplies timbre and recording conditions; the text supplies the words.
CosyVoice and Spark-TTS refine the recipe with a stronger text encoder, a supervised semantic stream, and flow-matching decoders, but they keep the core: a language model over codec tokens conditioned on text and a speaker prompt. The remaining weaknesses are the ones you would predict from that framing. Autoregressive token generation is slow enough to matter for real-time use, the models can hallucinate or drop words, and cloning raises the obvious consent and provenance questions taken up in the last part of this chapter.
Two-stage and streaming
Autoregressive structure, diffusion detail
The two-stage design splits the job. A first stage runs an autoregressive or flow-matching model that produces a compact, coarse token sequence — enough to fix the content and the rhythm. A second stage, often a diffusion transformer, takes that sequence and refines it into the fine acoustic tokens that the vocoder decodes. Systems such as GLM-TTS and DiTAR use this division so that the expensive autoregressive part generates few tokens and the detail is filled in by a parallel, non-autoregressive stage that can be run with few steps. The design pays off in streaming, where what matters is not total synthesis time but the latency of the first chunk.
For a conversational agent the number that users feel is the gap between the end of their speech and the first sound of the reply. That gap is the sum of the text-encoding time, the time to generate the first autoregressive chunk, and the time for the vocoder to render it; the rest of the utterance is synthesised while the first part is already playing. The budget bar makes the arithmetic concrete. The target for a natural turn is a few hundred milliseconds at most, and every stage in the chain gets a slice of it.
First-chunk latency budget, drawn with Guide.drawStacked from GenMedia.timeline.budget. The dashed line is a 300 ms target.
Where this shows up
Synthesis is half of every voice agent
The speaking half of a duplex loop
A voice agent pairs synthesis with recognition and turn-taking. The first-chunk budget here and the barge-in behaviour there are two halves of the same conversational contract, and the codec vocabulary is shared between them.
Everything rests on the codec
The token vocabulary, frame rate and bitrate come from the neural codec of the previous part. Change the codec and every latency number in this part changes with it.
Cloning and prosody are the two properties users judge first. Zero-shot cloning is now a commodity, which pushes the interesting questions to consent, watermarking and attribution — treated in the final part — while prosody, the right intonation for a question or an aside, remains the hardest thing to specify and to evaluate.
Further reading
The papers below mark the shift from aligned spectrogram prediction to codec language modelling. If you take away one thing, take away the alignment constraint: monotonicity is what separates speech synthesis from generic sequence transduction.
- Wang, Skerry-Ryan, Stanton and coauthors, "Tacotron: Towards End-to-End Speech Synthesis", 2017 — attention alignment learned without labels.
- Ren, Hu, Tan and coauthors, "FastSpeech: Fast, Robust and Controllable Text to Speech", 2019 — the explicit duration model.
- Wang, Chen, Wu and coauthors, "Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers", 2023 — VALL-E and zero-shot cloning.
- Du, Chen, Zhang and coauthors, "CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer", 2024 — semantic tokens and flow-matching decoding.
- Peng, Li, Zhang and coauthors, "Spark-TTS", 2025 — a codec-language-model synthesiser with controllable attributes.
Cheat sheet
| Term | Meaning here |
|---|---|
| Alignment | The correspondence between text symbols and audio frames |
| Monotonicity | The band moves forward only; speech never speaks a symbol twice out of order |
| Attention alignment | Tacotron learns the map as a soft by-product of frame prediction |
| Duration model | Explicit frames-per-symbol prediction; FastSpeech-style, fast and controllable |
| Codec language model | Text-conditioned prediction over acoustic tokens; VALL-E and relatives |
| Zero-shot cloning | Copying a voice from a short prompt with no fine-tuning |
| Two-stage synthesis | Coarse autoregressive tokens, then diffusion-refined detail |
| First-chunk latency | Text encoder + first AR chunk + vocoder; the gap a listener feels |
| Vocoder | The decoder from sound tokens or spectrograms to a waveform |