Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

Why a codec

A learned front end, not a filterbank

A waveform language model that predicts one sample at a time is doomed by the numbers, not by any modelling subtlety: at 24 kHz a ten-second clip is 240,000 tokens, an order of magnitude beyond a long context window, and every one of them is a near-continuous value with almost no redundancy at the sample level. Classical compression solves the redundancy problem but yields a continuous code that is awkward to model as a language. A neural codec solves both: it learns a compact representation and then quantises it so the output is a sequence of integers drawn from a small, fixed vocabulary.

The design is an autoencoder with a bottleneck that is deliberately discrete. An encoder — usually a stack of strided convolutions or a transformer over the waveform — reduces the 24 kHz signal to a sequence of latent vectors at some much lower frame rate, typically seventy-five vectors per second. A vector quantiser replaces each vector with indices into a codebook. A decoder mirrors the encoder and reconstructs the waveform. The reconstruction is not exact, and that is fine: the model only needs to preserve what a listener can hear.

The mel spectrogram the encoder effectively has to explain, computed with GenMedia.dsp.melSpectrogram. Compression is judged on how faithfully a decoder reproduces this picture, not the raw samples.

Read the heatmap top to bottom as frequency and left to right as time. Each vertical slice becomes one or a few codec frames; the frame rate decides how many slices per second are kept.

💡 By the end of this part you'll be able to describe the encoder, residual quantiser and decoder of a neural codec, explain how a residual stack of codebooks trades bitrate for fidelity, compute the token rate and bitrate from the frame rate and codebook count, and say when a codec should emit semantic rather than acoustic tokens.
2

The encoder and the residual quantiser

A stack of codebooks, each fixing the last one's error

One codebook is not enough. Quantising a continuous frame to the nearest of a few thousand entries leaves a residual error, and simply enlarging the codebook to cover the space runs into a lookup and training cost that grows badly. Residual vector quantisation takes the other route: quantise the frame with the first codebook, compute what is left over, quantise that with a second codebook, and repeat. Each stage sees only the error the previous stages could not represent, so the reconstruction improves monotonically as stages are added. The price is one token per stage per frame, which is exactly the trade the next step quantifies.

The demo below applies the quantiser directly to a short waveform rather than to encoder latents, so you can see the mechanism without a network in the way. A real codec runs the same arithmetic on a vector of latent features, but the residual structure — each stage chipping away at what remains — is identical. The dashed curve is the reconstruction after the current number of stages.

Original waveform (solid) and the residual-quantised reconstruction (dashed), both from GenMedia.dsp.rvq. Drag the stages slider and watch the dashed curve snap onto the solid one.

⚠ More stages is not free in a language model. Every extra codebook multiplies the token rate and therefore the sequence length the generator must model. Adaptive codecs spend stages where the signal is hard and skip them where it is easy, which is why an entropy-coded residual stack can reach the same quality at a lower average rate.
3

Frame rate, codebooks and bitrate

Three knobs, one budget

The token rate is the frame rate times the number of codebooks, and the bitrate adds the bits per codebook entry. Those are the three knobs and they are not independent: raise the frame rate and you resolve transients better but spend more tokens per second; add codebooks and you sharpen each frame but spend more tokens too; widen each codebook and the rate climbs without more frames. The reference point for speech is the EnCodec-style setting of seventy-five frames per second, eight residual codebooks and ten bits each, which lands at roughly six kilobits per second — a few hundred tokens per second, small enough for a language model to generate.

Below, the bars show the residual energy after each stage, in decibels relative to the un-quantised signal. The first stage removes the bulk of the error; later stages remove successively less, which is the point of adaptive allocation. Move the frame-rate and codebook sliders to watch the token and bitrate readout change; move the stages slider to watch the reconstruction error fall while the token count stands still.

Residual energy after each RVQ stage, in dB below the input variance. Bars are labelled with the stage; the first is the input itself at 0 dB.

💡 The arithmetic to remember. Token rate $= \text{frame rate} \times \text{codebooks}$. Bitrate $= \text{frame rate} \times \text{codebooks} \times \text{bits}$ bits per second. At 75 Hz, eight books and ten bits this is 600 tokens per second and 6 kbps; halving the frame rate halves both, and that is the single most effective lever on the sequence length a generator has to model.
4

Semantic vs acoustic tokens

Two vocabularies, two jobs

A single codec gives a token stream that is faithful to the waveform but carries no linguistic structure: neighbouring tokens are acoustic details, and a language model trained on them learns to imitate sound without necessarily learning to say anything. That is why many systems use two streams. Semantic tokens come from a representation distilled from a self-supervised speech model, so they encode phonetic content and vary slowly; acoustic tokens come from the residual codec and carry the fine detail — timbre, prosody and the exact waveform. A typical pipeline models the semantic stream with a language model and conditions an acoustic stage, or a vocoder, on it.

Once sound is a token stream, every technique from the language-model world applies: the sequence can be generated autoregressively, masked, or diffused, and the same vocabulary can be shared across tasks. Speech recognition, synthesis, translation and spoken dialogue can all be framed as prediction over codec tokens, which is why the codec is the hinge of this chapter rather than a preprocessing detail.

5

Where this shows up

The codec is the common interface

Sound · Synthesis

Text-to-speech over codec tokens

A codec language model treats the acoustic token stream as text to be predicted, and a vocoder or diffusion decoder turns the predicted stream back into a waveform. The frame rate chosen here sets how much audio one generated token covers.

Sound · Recognition

Discrete targets for recognition

Speech-to-text models can emit codec or semantic tokens directly, which unifies recognition and synthesis in one vocabulary. That is what makes a single model able to both transcribe and speak in a full-duplex conversation.

The next part builds on this vocabulary directly. With sound reduced to a few hundred tokens per second, a text-to-speech system becomes a conditional language model over that vocabulary plus a decoder, and the interesting engineering moves from signal processing to alignment, controllability and latency.

Further reading

These papers trace the codec from the first neural audio compressor to the multi-codebook designs that modern speech models use. If you take away one thing, take away the residual stack: it is why a codec can be small and faithful at the same time.

Cheat sheet

TermMeaning here
EncoderTurns the waveform into a sequence of latent vectors at the frame rate
Frame rateLatent vectors per second; typically 50–75 for speech
Vector quantiserReplaces a latent vector with indices into a codebook
Residual quantisationEach codebook quantises what the previous stages left over
CodebookThe finite set of entries one stage chooses among
Bits per codebookCodebook size: 10 bits is 1024 entries
Token rateframe rate × codebooks per second
Bitratetoken rate × bits per token; ~6 kbps at the reference setting
Semantic tokensSlow, phonetic codes distilled from a self-supervised model
Acoustic tokensFine-grained codec codes carrying timbre and detail
VocoderThe decoder that turns codes or spectrograms back into a waveform
7

Check your understanding

0/4 answered