Neural audio codecs
The previous part left a twenty-four-thousand-token-per-second waveform on the table. A neural audio codec is the learned front end that replaces it with a few hundred discrete tokens per second, and it does so in three pieces: an encoder that turns the waveform into a short sequence of continuous vectors, a residual vector quantiser that snaps each vector to a stack of codebook entries, and a decoder that turns the codes back into sound. The codes are the interface every modern speech and audio model speaks, so this part is really about the vocabulary of the sound chapter: how big it is, how fast it ticks, and what the frame rate and bitrate knobs actually buy.
Why a codec
A learned front end, not a filterbank
A waveform language model that predicts one sample at a time is doomed by the numbers, not by any modelling subtlety: at 24 kHz a ten-second clip is 240,000 tokens, an order of magnitude beyond a long context window, and every one of them is a near-continuous value with almost no redundancy at the sample level. Classical compression solves the redundancy problem but yields a continuous code that is awkward to model as a language. A neural codec solves both: it learns a compact representation and then quantises it so the output is a sequence of integers drawn from a small, fixed vocabulary.
The design is an autoencoder with a bottleneck that is deliberately discrete. An encoder — usually a stack of strided convolutions or a transformer over the waveform — reduces the 24 kHz signal to a sequence of latent vectors at some much lower frame rate, typically seventy-five vectors per second. A vector quantiser replaces each vector with indices into a codebook. A decoder mirrors the encoder and reconstructs the waveform. The reconstruction is not exact, and that is fine: the model only needs to preserve what a listener can hear.
The mel spectrogram the encoder effectively has to explain, computed with GenMedia.dsp.melSpectrogram. Compression is judged on how faithfully a decoder reproduces this picture, not the raw samples.
Read the heatmap top to bottom as frequency and left to right as time. Each vertical slice becomes one or a few codec frames; the frame rate decides how many slices per second are kept.
The encoder and the residual quantiser
A stack of codebooks, each fixing the last one's error
One codebook is not enough. Quantising a continuous frame to the nearest of a few thousand entries leaves a residual error, and simply enlarging the codebook to cover the space runs into a lookup and training cost that grows badly. Residual vector quantisation takes the other route: quantise the frame with the first codebook, compute what is left over, quantise that with a second codebook, and repeat. Each stage sees only the error the previous stages could not represent, so the reconstruction improves monotonically as stages are added. The price is one token per stage per frame, which is exactly the trade the next step quantifies.
The demo below applies the quantiser directly to a short waveform rather than to encoder latents, so you can see the mechanism without a network in the way. A real codec runs the same arithmetic on a vector of latent features, but the residual structure — each stage chipping away at what remains — is identical. The dashed curve is the reconstruction after the current number of stages.
Original waveform (solid) and the residual-quantised reconstruction (dashed), both from GenMedia.dsp.rvq. Drag the stages slider and watch the dashed curve snap onto the solid one.
Frame rate, codebooks and bitrate
Three knobs, one budget
The token rate is the frame rate times the number of codebooks, and the bitrate adds the bits per codebook entry. Those are the three knobs and they are not independent: raise the frame rate and you resolve transients better but spend more tokens per second; add codebooks and you sharpen each frame but spend more tokens too; widen each codebook and the rate climbs without more frames. The reference point for speech is the EnCodec-style setting of seventy-five frames per second, eight residual codebooks and ten bits each, which lands at roughly six kilobits per second — a few hundred tokens per second, small enough for a language model to generate.
Below, the bars show the residual energy after each stage, in decibels relative to the un-quantised signal. The first stage removes the bulk of the error; later stages remove successively less, which is the point of adaptive allocation. Move the frame-rate and codebook sliders to watch the token and bitrate readout change; move the stages slider to watch the reconstruction error fall while the token count stands still.
Residual energy after each RVQ stage, in dB below the input variance. Bars are labelled with the stage; the first is the input itself at 0 dB.
Semantic vs acoustic tokens
Two vocabularies, two jobs
A single codec gives a token stream that is faithful to the waveform but carries no linguistic structure: neighbouring tokens are acoustic details, and a language model trained on them learns to imitate sound without necessarily learning to say anything. That is why many systems use two streams. Semantic tokens come from a representation distilled from a self-supervised speech model, so they encode phonetic content and vary slowly; acoustic tokens come from the residual codec and carry the fine detail — timbre, prosody and the exact waveform. A typical pipeline models the semantic stream with a language model and conditions an acoustic stage, or a vocoder, on it.
Once sound is a token stream, every technique from the language-model world applies: the sequence can be generated autoregressively, masked, or diffused, and the same vocabulary can be shared across tasks. Speech recognition, synthesis, translation and spoken dialogue can all be framed as prediction over codec tokens, which is why the codec is the hinge of this chapter rather than a preprocessing detail.
Where this shows up
The codec is the common interface
Text-to-speech over codec tokens
A codec language model treats the acoustic token stream as text to be predicted, and a vocoder or diffusion decoder turns the predicted stream back into a waveform. The frame rate chosen here sets how much audio one generated token covers.
Discrete targets for recognition
Speech-to-text models can emit codec or semantic tokens directly, which unifies recognition and synthesis in one vocabulary. That is what makes a single model able to both transcribe and speak in a full-duplex conversation.
The next part builds on this vocabulary directly. With sound reduced to a few hundred tokens per second, a text-to-speech system becomes a conditional language model over that vocabulary plus a decoder, and the interesting engineering moves from signal processing to alignment, controllability and latency.
Further reading
These papers trace the codec from the first neural audio compressor to the multi-codebook designs that modern speech models use. If you take away one thing, take away the residual stack: it is why a codec can be small and faithful at the same time.
- Défossez, Copet, Synnaeve and Adi, "High Fidelity Neural Audio Compression", 2022 — EnCodec and the 75 Hz, multi-codebook setting quoted throughout.
- Zeghidour, Luebs, Omran and coauthors, "SoundStream: An End-to-End Neural Audio Codec", 2021 — residual vector quantisation and the quantiser-dropout trick.
- Borsos, Marinier, Vincent and coauthors, "AudioLM: A Language Modeling Approach to Audio Generation", 2022 — semantic and acoustic tokens in one pipeline.
- Kumar, Seetharaman, Luebs and coauthors, "High-Fidelity Audio Compression with Improved RVQGAN", 2023 — an improved residual quantiser and discriminator design.
Cheat sheet
| Term | Meaning here |
|---|---|
| Encoder | Turns the waveform into a sequence of latent vectors at the frame rate |
| Frame rate | Latent vectors per second; typically 50–75 for speech |
| Vector quantiser | Replaces a latent vector with indices into a codebook |
| Residual quantisation | Each codebook quantises what the previous stages left over |
| Codebook | The finite set of entries one stage chooses among |
| Bits per codebook | Codebook size: 10 bits is 1024 entries |
| Token rate | frame rate × codebooks per second |
| Bitrate | token rate × bits per token; ~6 kbps at the reference setting |
| Semantic tokens | Slow, phonetic codes distilled from a self-supervised model |
| Acoustic tokens | Fine-grained codec codes carrying timbre and detail |
| Vocoder | The decoder that turns codes or spectrograms back into a waveform |