What audio is, numerically
A sound is a pressure wave, and a microphone turns it into a list of numbers. That is the whole of the representation problem: the wave is one-dimensional, so it seems simpler than an image, but the list is enormous. At twenty-four kilohertz a single second is twenty-four thousand floats, and a model that emits them one at a time is emitting twenty-four thousand tokens for one second of sound. This part builds the picture the rest of the sound chapter depends on: the waveform, the sample rate and the Nyquist limit, the short-time Fourier transform with its window-and-hop trade, the mel scale that compresses frequency by ear rather than by hertz, and the token arithmetic that makes a raw waveform language model untenable.
A sound is a list of numbers
One axis, sampled fast
A waveform is a function of time: at each instant the microphone reports the air pressure, positive or negative relative to the ambient level. To store it on a computer you sample that function at regular intervals and keep the values. The number of samples per second is the sample rate, written $f_s$, and the time between samples is $1/f_s$. Speech telephony uses eight kilohertz, wideband speech sixteen, music and high-fidelity speech forty-four point one or forty-eight, and the models in this chapter usually work at twenty-four. The demo below draws a sum of three tones — a fundamental and two partials, which is what makes it read as a chord rather than a beep — and lets you change its pitch and its length.
Nothing about the picture is digital yet: the list is a faithful-enough approximation to a continuous wave, provided we sampled fast enough, which is the subject of the next step. What matters here is only that the signal is an array of scalar values with a single index. A colormap of it would be a one-pixel-wide image; a spectrogram of it will be a two-dimensional picture, and that is the representation every audio model actually consumes.
A chord of three tones sampled at 24 kHz and drawn with Guide.drawLines. Amplitude is normalised to the range −1 to 1.
Sampling and the spectrum
The one theorem you cannot skip
Sampling at $f_s$ can only represent frequencies up to the Nyquist frequency $f_s/2$. A component above that limit is not lost cleanly; it folds back and appears as a lower frequency, an artefact called aliasing, the same way a wagon wheel on film can appear to spin backwards. This is why a real pipeline low-pass filters before it downsamples. It is also why the sample rate is a hard ceiling on what a model can ever represent: at 24 kHz nothing above 12 kHz exists in the data, and most of the perceptual detail above that ceiling is inaudible to adults anyway.
The spectrum is the other view of the same signal. A discrete Fourier transform of one window of samples reports, for each frequency bin, how much of that frequency is present. The canvas below takes a single short window of the chord and plots its magnitude against frequency, with a dashed marker at the fundamental. The peaks are the partials; the rest is leakage and the tail of the window. Almost every idea in audio generation is a statement about this picture rather than about the waveform, because the spectrum is where the structure lives.
Magnitude spectrum of one window, from GenMedia.dsp.stft. The dashed line marks the fundamental.
The right edge is the Nyquist frequency, half the sample rate. Everything a model can learn about this sound is summarised by peaks like these, tracked frame by frame.
Windowing and the STFT
A trade you cannot avoid
A single Fourier transform of the whole signal tells you which frequencies are present but not when. The short-time Fourier transform fixes that by transforming overlapping short windows and stacking the results: each window gives one column of a two-dimensional picture whose axes are time and frequency. The window is not rectangular — a smooth taper such as a Hann window reduces the spectral leakage that a hard cut would create — and the amount of overlap is the hop.
Window length sets a trade with no free side. A long window resolves frequency finely but smears time, so a plosive or a note onset is blurred; a short window pins down the moment but blurs neighbouring frequencies. The hop sets the frame rate independently: hop equals window over two is the usual starting point, and smaller hops oversample time at the cost of more frames. The readout tracks both, and you can drag the sliders to watch the number of frames and the frequency resolution move in opposite directions.
The same spectrum at the chosen window and hop, redrawn with Guide.drawLines: a short window broadens every peak, a long window sharpens it and smears the frame in time.
Frequency resolution is $f_s/\text{win}$ hertz per bin; time resolution is $\text{hop}/f_s$ milliseconds per frame. Their product is fixed by the sampling, which is the uncertainty principle wearing engineering clothes.
Mel and what it throws away
Compress frequency the way an ear does
Hertz are linear, but hearing is not. We resolve a few hertz near the bottom of the range and hundreds of hertz at the top, so a linear frequency axis spends most of its bins on detail nobody can hear. The mel scale rescales frequency so that equal distances sound equal: it is roughly linear below a kilohertz and logarithmic above. A mel filterbank is a bank of overlapping triangles spaced evenly on that scale, and applying it collapses hundreds of FFT bins into a few dozen bands. The heatmap below is the result — a mel spectrogram, the standard input to a speech recogniser and a common intermediate for a synthesiser.
What it throws away is resolution in the high frequencies and all phase. For recognition that is fine, indeed helpful, because it matches the ear and shrinks the input. For generation it is a loss you have to undo: a vocoder or a diffusion decoder has to invent the fine structure and the phase back from the compressed representation, which is exactly why the neural codec in the next part keeps a learned latent rather than a plain mel spectrogram.
Mel spectrogram of the chord, computed with GenMedia.dsp.melSpectrogram and drawn as a heatmap. Darker means more energy.
Compare with the linear spectrum above: the low partials still show, but the high region is pooled into a handful of wide bands.
Where this shows up
Every audio model starts from these four pictures
From samples to tokens
The neural codec is a learned front end that replaces the mel filterbank. It maps a waveform to a low-rate sequence of discrete codes and back, and it keeps the fine structure the mel scale discards, which is why it can drive a vocoder convincingly.
Alignment sits on the spectrogram
Text-to-speech alignment and duration models and speech-to-text CTC and RNN-T alignment are both statements about how a sequence of text symbols lines up with the frames of a spectrogram. The window and hop chosen here set the frame rate both of them operate at.
The same machinery reappears outside speech. Music generation, audio editing and environmental-sound models all consume a spectrogram or a learned codec latent, and all of them are constrained by the sample rate and the STFT trade you can see above. The next part takes the one step this one left out: how a network learns its own front end instead of using a fixed filterbank.
Further reading
The sources below cover the classical signal processing first and the learned front ends second. If you take away one idea, take away the token-rate arithmetic: it explains why audio generation is a codec problem before it is a language-model problem.
Smith's Mathematics of the Discrete Fourier Transform is the free, careful treatment of the transform; Oppenheim and Schafer is the standard reference for the STFT and windowing; and the mel-scale original is worth reading once for the motivation rather than the constants. The EnCodec paper is the bridge into the next part.
- Julius O. Smith III, Mathematics of the Discrete Fourier Transform, online book — the DFT, the window and the leakage it controls.
- Alan Oppenheim and Ronald Schafer, Discrete-Time Signal Processing — sampling, aliasing and the short-time Fourier transform.
- Stevens, Volkmann and Newman, "A Scale for the Psychological Magnitude of Pitch", 1937 — the mel scale on which the filterbank is built.
- Défossez, Copet, Synnaeve and Adi, "High Fidelity Neural Audio Compression", 2022 — EnCodec, the codec of the next part.
Cheat sheet
| Term | Meaning here |
|---|---|
| Waveform | The signal as amplitude against time; one scalar per sample |
| Sample rate $f_s$ | Samples per second; 24 kHz for the models in this chapter |
| Nyquist frequency | $f_s/2$; the highest frequency the samples can represent |
| Aliasing | Energy above Nyquist folding back to a false lower frequency |
| Spectrum | Magnitude per frequency bin, from a DFT of one window |
| STFT | Overlapping windowed transforms stacked into time × frequency |
| Window length | Trades frequency resolution against time resolution |
| Hop | Samples between windows; sets the frame rate |
| Mel scale | Perceptual frequency, roughly linear then logarithmic |
| Mel spectrogram | Filterbank output: a few dozen bands, phase discarded |
| Raw token rate | $f_s$ tokens per second at one token per sample |
| Codec token rate | frame rate × codebooks; 75 × 8 = 600 per second |