Reading and display settings

Appearance

System follows your operating system and keeps following it, even if you change it later. The header's sun, moon and monitor cycle the same three options.

Text size (%) 100%

Default. Scales every text size on the site, equations and tables included.

Reading width 70ch

How much text runs across one line of prose. Narrower is easier to track; wider fits more on screen.

Line spacing 1.6

The leading on body text. Taller leading helps a tired eye stay on the line.

Density

Padding and gaps around controls, cards, and tables — how much breathing room the layout leaves itself.

Motion

System follows your operating system. Reduced removes every transition on this site. Full keeps them on unless your system asks for less.

1

A sound is a list of numbers

One axis, sampled fast

A waveform is a function of time: at each instant the microphone reports the air pressure, positive or negative relative to the ambient level. To store it on a computer you sample that function at regular intervals and keep the values. The number of samples per second is the sample rate, written $f_s$, and the time between samples is $1/f_s$. Speech telephony uses eight kilohertz, wideband speech sixteen, music and high-fidelity speech forty-four point one or forty-eight, and the models in this chapter usually work at twenty-four. The demo below draws a sum of three tones — a fundamental and two partials, which is what makes it read as a chord rather than a beep — and lets you change its pitch and its length.

Nothing about the picture is digital yet: the list is a faithful-enough approximation to a continuous wave, provided we sampled fast enough, which is the subject of the next step. What matters here is only that the signal is an array of scalar values with a single index. A colormap of it would be a one-pixel-wide image; a spectrogram of it will be a two-dimensional picture, and that is the representation every audio model actually consumes.

A chord of three tones sampled at 24 kHz and drawn with Guide.drawLines. Amplitude is normalised to the range −1 to 1.

💡 By the end of this part you'll be able to say what a waveform, a sample rate and a Nyquist limit are, read a spectrogram and a mel spectrogram, choose a window and hop for the STFT trade, and compute the token rate that a raw-waveform model would need against a codec's.
2

Sampling and the spectrum

The one theorem you cannot skip

Sampling at $f_s$ can only represent frequencies up to the Nyquist frequency $f_s/2$. A component above that limit is not lost cleanly; it folds back and appears as a lower frequency, an artefact called aliasing, the same way a wagon wheel on film can appear to spin backwards. This is why a real pipeline low-pass filters before it downsamples. It is also why the sample rate is a hard ceiling on what a model can ever represent: at 24 kHz nothing above 12 kHz exists in the data, and most of the perceptual detail above that ceiling is inaudible to adults anyway.

The spectrum is the other view of the same signal. A discrete Fourier transform of one window of samples reports, for each frequency bin, how much of that frequency is present. The canvas below takes a single short window of the chord and plots its magnitude against frequency, with a dashed marker at the fundamental. The peaks are the partials; the rest is leakage and the tail of the window. Almost every idea in audio generation is a statement about this picture rather than about the waveform, because the spectrum is where the structure lives.

Magnitude spectrum of one window, from GenMedia.dsp.stft. The dashed line marks the fundamental.

The right edge is the Nyquist frequency, half the sample rate. Everything a model can learn about this sound is summarised by peaks like these, tracked frame by frame.

⚠ Sample rate is not quality. Raising $f_s$ buys bandwidth, not fidelity for its own sake; a 48 kHz recording of a muffled source still contains nothing above a few kilohertz. The reason the models here use 24 kHz is that it is the smallest rate that keeps speech and most music sounding natural, which keeps the raw token count as small as it can be before any codec.
3

Windowing and the STFT

A trade you cannot avoid

A single Fourier transform of the whole signal tells you which frequencies are present but not when. The short-time Fourier transform fixes that by transforming overlapping short windows and stacking the results: each window gives one column of a two-dimensional picture whose axes are time and frequency. The window is not rectangular — a smooth taper such as a Hann window reduces the spectral leakage that a hard cut would create — and the amount of overlap is the hop.

Window length sets a trade with no free side. A long window resolves frequency finely but smears time, so a plosive or a note onset is blurred; a short window pins down the moment but blurs neighbouring frequencies. The hop sets the frame rate independently: hop equals window over two is the usual starting point, and smaller hops oversample time at the cost of more frames. The readout tracks both, and you can drag the sliders to watch the number of frames and the frequency resolution move in opposite directions.

The same spectrum at the chosen window and hop, redrawn with Guide.drawLines: a short window broadens every peak, a long window sharpens it and smears the frame in time.

Frequency resolution is $f_s/\text{win}$ hertz per bin; time resolution is $\text{hop}/f_s$ milliseconds per frame. Their product is fixed by the sampling, which is the uncertainty principle wearing engineering clothes.

4

Mel and what it throws away

Compress frequency the way an ear does

Hertz are linear, but hearing is not. We resolve a few hertz near the bottom of the range and hundreds of hertz at the top, so a linear frequency axis spends most of its bins on detail nobody can hear. The mel scale rescales frequency so that equal distances sound equal: it is roughly linear below a kilohertz and logarithmic above. A mel filterbank is a bank of overlapping triangles spaced evenly on that scale, and applying it collapses hundreds of FFT bins into a few dozen bands. The heatmap below is the result — a mel spectrogram, the standard input to a speech recogniser and a common intermediate for a synthesiser.

What it throws away is resolution in the high frequencies and all phase. For recognition that is fine, indeed helpful, because it matches the ear and shrinks the input. For generation it is a loss you have to undo: a vocoder or a diffusion decoder has to invent the fine structure and the phase back from the compressed representation, which is exactly why the neural codec in the next part keeps a learned latent rather than a plain mel spectrogram.

Mel spectrogram of the chord, computed with GenMedia.dsp.melSpectrogram and drawn as a heatmap. Darker means more energy.

Compare with the linear spectrum above: the low partials still show, but the high region is pooled into a handful of wide bands.

💡 The token-rate argument. A raw model at 24 kHz emits one token per sample — twenty-four thousand tokens per second. A neural codec at 75 frames per second with eight residual codebooks emits 75 × 8 = 600 tokens per second of sound, forty times fewer. That factor is the entire reason the rest of this chapter speaks in codec tokens, not samples.
5

Where this shows up

Every audio model starts from these four pictures

Sound · Codecs

From samples to tokens

The neural codec is a learned front end that replaces the mel filterbank. It maps a waveform to a low-rate sequence of discrete codes and back, and it keeps the fine structure the mel scale discards, which is why it can drive a vocoder convincingly.

Sound · Speech

Alignment sits on the spectrogram

Text-to-speech alignment and duration models and speech-to-text CTC and RNN-T alignment are both statements about how a sequence of text symbols lines up with the frames of a spectrogram. The window and hop chosen here set the frame rate both of them operate at.

The same machinery reappears outside speech. Music generation, audio editing and environmental-sound models all consume a spectrogram or a learned codec latent, and all of them are constrained by the sample rate and the STFT trade you can see above. The next part takes the one step this one left out: how a network learns its own front end instead of using a fixed filterbank.

Further reading

The sources below cover the classical signal processing first and the learned front ends second. If you take away one idea, take away the token-rate arithmetic: it explains why audio generation is a codec problem before it is a language-model problem.

Smith's Mathematics of the Discrete Fourier Transform is the free, careful treatment of the transform; Oppenheim and Schafer is the standard reference for the STFT and windowing; and the mel-scale original is worth reading once for the motivation rather than the constants. The EnCodec paper is the bridge into the next part.

Cheat sheet

TermMeaning here
WaveformThe signal as amplitude against time; one scalar per sample
Sample rate $f_s$Samples per second; 24 kHz for the models in this chapter
Nyquist frequency$f_s/2$; the highest frequency the samples can represent
AliasingEnergy above Nyquist folding back to a false lower frequency
SpectrumMagnitude per frequency bin, from a DFT of one window
STFTOverlapping windowed transforms stacked into time × frequency
Window lengthTrades frequency resolution against time resolution
HopSamples between windows; sets the frame rate
Mel scalePerceptual frequency, roughly linear then logarithmic
Mel spectrogramFilterbank output: a few dozen bands, phase discarded
Raw token rate$f_s$ tokens per second at one token per sample
Codec token rateframe rate × codebooks; 75 × 8 = 600 per second
7

Check your understanding

0/4 answered