Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

From text to sound

Five stages, and one alignment in the middle

The pipeline below is the shape almost every modern system shares. Text is normalised and turned into symbols. An encoder gives each symbol a vector. An alignment stage expands those vectors along the time axis, deciding how many frames of sound each symbol owns. An acoustic model produces a representation of sound — historically a mel spectrogram, now usually codec tokens — and a vocoder or decoder turns that representation into a waveform. The alignment stage is the one that has no counterpart in machine translation, and it is where the design battles of the last decade were fought.

The reason alignment is hard is that it is a many-to-many correspondence with no labels at training time. We know the text and we know the audio, and we do not know which frames belong to which phoneme. Early systems learned it implicitly with attention, which discovers a soft monotonic map as a side effect of predicting spectrogram frames. Later systems predicted it explicitly with a duration model, which is faster and more controllable. The codec language models of the last few years sidestep explicit alignment almost entirely by treating speech as a token sequence and letting a strong language model learn the correspondence from data.

The text-to-speech pipeline. The middle box — alignment and duration — is the stage the rest of this part is about.

Read the diagram left to right. Text becomes vectors, vectors are stretched in time, the stretched sequence becomes sound tokens, and the tokens become a waveform.

💡 By the end of this part you'll be able to explain why alignment needs monotonicity, describe attention-era alignment against explicit duration models, outline a codec language model such as VALL-E, and account for the first-chunk latency of a streaming two-stage synthesiser.
2

Alignment and duration

Which frames belong to which symbol

Attention-based synthesisers such as Tacotron let the decoder attend over the encoder's symbol vectors while generating frames one at a time. Trained on enough data, the attention settles into a roughly diagonal band: each frame looks at the symbol currently being spoken, and the band moves left to right as the utterance progresses. That band is the alignment, learned without labels. It is also fragile. Because the map is soft, the decoder can skip a symbol, repeat it, or stall, and failures are catastrophic rather than gradual.

Explicit duration models replaced the soft map with a hard one. A duration predictor estimates, for each symbol, how many frames it spans; at inference those counts are used to expand the symbol sequence to frame length, and the acoustic model becomes a non-autoregressive mapping from the enlarged sequence to spectrograms. This is the design of FastSpeech and its descendants. It is fast, it parallelises, and it makes timing directly controllable — you can slow a word down by editing its duration. The cost is that a duration error becomes an audible timing error rather than something attention can smooth away.

The heatmap is a hard alignment of the kind a duration model produces: a bright diagonal band, one frame column per time step, one row per symbol. Move the sliders to change how many symbols there are and how long the utterance runs; the band steepens or flattens but stays monotonic, because speech does not go backwards.

A monotonic alignment matrix: rows are text symbols, columns are frames. Bright means the symbol is being spoken in that frame.

⚠ Monotonicity is a constraint, not a hope. Nothing about a plain attention layer forbids looking backwards. Every attention synthesiser relies on either the data or an explicit monotonic bias to keep the band from folding, and when it folds you hear a skipped or repeated word. Explicit durations are what made the failure mode disappear.
3

Codec language models

Speech as text

Once a codec has reduced speech to a token stream, synthesis becomes conditional language modelling: given the text tokens, predict the acoustic tokens. VALL-E framed it exactly this way and added the property that made the field take notice — zero-shot cloning. Because the model is conditioned on both the text and a short acoustic prompt, three seconds of a target speaker's audio are enough to make it continue in that voice, with no fine-tuning and no per-speaker parameters. The prompt supplies timbre and recording conditions; the text supplies the words.

CosyVoice and Spark-TTS refine the recipe with a stronger text encoder, a supervised semantic stream, and flow-matching decoders, but they keep the core: a language model over codec tokens conditioned on text and a speaker prompt. The remaining weaknesses are the ones you would predict from that framing. Autoregressive token generation is slow enough to matter for real-time use, the models can hallucinate or drop words, and cloning raises the obvious consent and provenance questions taken up in the last part of this chapter.

4

Two-stage and streaming

Autoregressive structure, diffusion detail

The two-stage design splits the job. A first stage runs an autoregressive or flow-matching model that produces a compact, coarse token sequence — enough to fix the content and the rhythm. A second stage, often a diffusion transformer, takes that sequence and refines it into the fine acoustic tokens that the vocoder decodes. Systems such as GLM-TTS and DiTAR use this division so that the expensive autoregressive part generates few tokens and the detail is filled in by a parallel, non-autoregressive stage that can be run with few steps. The design pays off in streaming, where what matters is not total synthesis time but the latency of the first chunk.

For a conversational agent the number that users feel is the gap between the end of their speech and the first sound of the reply. That gap is the sum of the text-encoding time, the time to generate the first autoregressive chunk, and the time for the vocoder to render it; the rest of the utterance is synthesised while the first part is already playing. The budget bar makes the arithmetic concrete. The target for a natural turn is a few hundred milliseconds at most, and every stage in the chain gets a slice of it.

First-chunk latency budget, drawn with Guide.drawStacked from GenMedia.timeline.budget. The dashed line is a 300 ms target.

💡 Streaming changes the objective. A model that takes a full second to synthesise five seconds of speech can still feel instant if the first hundred milliseconds arrive within the budget. Streaming synthesis is therefore an exercise in producing an early, good-enough first chunk, not in making the whole utterance fast.
5

Where this shows up

Synthesis is half of every voice agent

Sound · Dialogue

The speaking half of a duplex loop

A voice agent pairs synthesis with recognition and turn-taking. The first-chunk budget here and the barge-in behaviour there are two halves of the same conversational contract, and the codec vocabulary is shared between them.

Sound · Foundations

Everything rests on the codec

The token vocabulary, frame rate and bitrate come from the neural codec of the previous part. Change the codec and every latency number in this part changes with it.

Cloning and prosody are the two properties users judge first. Zero-shot cloning is now a commodity, which pushes the interesting questions to consent, watermarking and attribution — treated in the final part — while prosody, the right intonation for a question or an aside, remains the hardest thing to specify and to evaluate.

Further reading

The papers below mark the shift from aligned spectrogram prediction to codec language modelling. If you take away one thing, take away the alignment constraint: monotonicity is what separates speech synthesis from generic sequence transduction.

Cheat sheet

TermMeaning here
AlignmentThe correspondence between text symbols and audio frames
MonotonicityThe band moves forward only; speech never speaks a symbol twice out of order
Attention alignmentTacotron learns the map as a soft by-product of frame prediction
Duration modelExplicit frames-per-symbol prediction; FastSpeech-style, fast and controllable
Codec language modelText-conditioned prediction over acoustic tokens; VALL-E and relatives
Zero-shot cloningCopying a voice from a short prompt with no fine-tuning
Two-stage synthesisCoarse autoregressive tokens, then diffusion-refined detail
First-chunk latencyText encoder + first AR chunk + vocoder; the gap a listener feels
VocoderThe decoder from sound tokens or spectrograms to a waveform
7

Check your understanding

0/4 answered