Listening models
A waveform — the raw stream of pressure measurements a microphone records, one number per instant of time — is the least token-shaped signal in the zoo. Every other modality arrives already broken into neat units: a sentence into words, an image into a grid of pixels or patches. A waveform is continuous instead. It has no natural frame boundaries, no obvious place to cut it, and the same sentence can be spoken in a whisper, a shout or a monotone without changing a word of it. A model that listens therefore has to choose between two front ends, a front end being whatever turns a raw signal into the sequence of tokens the rest of the model consumes. It can run a speech recogniser first — a model trained to convert sound into written words — and feed the trunk a transcript of plain text, which is cheap, readable, and deaf to how the words were said. Or it can learn a vocabulary of audio tokens, discrete symbols standing for short slices of sound, and put them straight into the trunk, so the model hears the sigh as well as the sentence. This part builds both routes and measures the difference.
Audio as input
A time-frequency picture
Nothing in a transformer wants a waveform. A transformer consumes a sequence of vectors of fixed width, and a waveform is one very long list of individual numbers with nothing to tell the model where one meaningful chunk ends and the next begins. The conventional first move is the one signal processing has used for decades, and it turns the sound into a picture. First, cut the samples — the individual amplitude measurements that make up the waveform — into short overlapping windows, spans of a few milliseconds considered together. Second, take a short-time Fourier transform of each window, which asks how much energy the signal carries at each frequency within that short span. Third, stack the resulting magnitude spectra side by side into a spectrogram, a two-dimensional picture with time along one axis, frequency along the other, and energy as brightness. Picture a contact sheet of narrow vertical tiles, each tile one window in time and its height read bottom to top from low to high frequency; a spectrogram is exactly that layout, and reading it left to right walks through the sound. Warp the frequency axis onto the mel scale, which matches how hearing allocates resolution — human ears separate low frequencies finely and high frequencies coarsely — and the result is a mel spectrogram. That picture is what audio encoders consume, an encoder being a network that reads a signal and produces a vector summary of it. Concretely, a mel spectrogram of a 16 kHz signal at a 25 ms window and a 10 ms hop is a hundred frames per second, where a frame is one thin vertical tile and a hop is the step from one window to the next. Each frame is a short vector of mel bands, so the whole spectrogram is a grid of numbers the encoder can treat much like a small image.
The sliders change the analysis window and the hop, and those two numbers decide the trade every spectrogram makes. A long window observes more of the signal at once, so it resolves frequency finely but smears time: two clicks close together blur into one event. A short window does the reverse, pinning down when something happened but leaving its pitch vague. The hop then sets how often you take a new reading, and therefore the number of frames per second of audio. Because each frame becomes one token, the hop alone sets the token count a native model will eventually pay, and a small hop is a finer dice over time — more, smaller pieces — bought at proportionally more context.
A mel spectrogram of a chirp plus a tone, computed with GenMedia.dsp.melSpectrogram and drawn with Guide.drawHeatmap. The rising band is the chirp; the straight line is the tone.
The encoder route
Hear words, then read them
The encoder route is the pipeline this site already treats in full elsewhere, and it is the one most people meet first. A trained speech recogniser — automatic speech recognition, or ASR, is the task of turning sound into written words — consumes the spectrogram and emits text. The trunk then receives that text as ordinary subwords, the word pieces a tokeniser splits text into. It is the cheapest possible interface to a language model, and the reason is the size of what crosses the boundary. The sequence it produces is tiny: a few subword tokens per second of speech, compared to the hundreds of frames per second the recogniser itself had to look at. The recogniser has already done the hard part, and it hands the trunk a short, tidy summary. The catch is that this interface is lossy in a specific and important way. A transcript preserves what was said and discards how it was said. Pitch, loudness, tempo, hesitations, laughter and the identity of the speaker are all squeezed out, because the recogniser was trained to be invariant to exactly those things: two recordings that differ only in delivery should produce the same words. That is a feature for transcription and a loss for anything that cares about the performance.
The diagram shows the route end to end. Everything to the left of the transcript is recogniser work: signal processing and a specialised acoustic model getting the words out. Everything to the right is ordinary language modelling, the same kind of next-token prediction the trunk does for any text. The interface between the two halves is a string of characters, which is why this design is sometimes called cascaded — two models in a chain rather than one model with a new sense. That seam is where the information is lost, and no amount of language modelling downstream can recover what the recogniser did not write down.
The encoder route: waveform, mel spectrogram, recogniser, transcript, trunk. The transcript is a handful of subword tokens per second; the spectrogram underneath was a hundred frames per second.
The recogniser is a strong, separately trained model. Using it as the interface means the trunk inherits its errors and its invariances, and never sees the signal it threw away.
The native-token route
Give the trunk a vocabulary for sound
The native-token route replaces the recogniser with a neural audio codec. A codec is a pair of networks that compress a signal down to a compact code and reconstruct it again, and here the compression half is the interesting one. An encoder maps the spectrogram to a small number of discrete audio tokens per frame, each token a symbol drawn from a fixed list, the same shape of object as a word piece in text, and a residual vector quantiser supplies the vocabulary those symbols come from. Residual vector quantisation, or RVQ, works in rounds: the first quantiser picks the closest match from its list and notes how much error is left over, the next quantiser encodes only that leftover error, and each further round cleans up what the previous ones missed. Each round has its own list, called a codebook, so with eight codebooks a frame is described by eight tokens instead of one, and a fine residual can be reconstructed without ever needing an enormous single list. The token rate is how many of these tokens the codec emits per second of audio, and it is a choice, not a property of the signal. A codec at 75 frames per second with eight residual codebooks emits 600 tokens per second; drop to 50 frames and four books and it is 200. The bitrate follows the same arithmetic, the token rate multiplied by the bits each token carries, and the codec's job is to keep the audio recognisable at whatever rate you pick: spend more tokens and the reconstruction gets closer to the original.
Those tokens go into the same trunk as the text, through an audio embedder — the small front-end network that maps each discrete audio token to a vector of the trunk's width, exactly as a text embedding table does for subwords. Putting the tokens into the trunk itself is what makes a model native rather than cascaded: there is no separate recogniser standing in front, and the audio and the words share one sequence, one attention mechanism and one context budget. The bars compare the token rates of the three front ends on the same axis: subword transcripts, mel frames and codec tokens. Native audio is one to two orders of magnitude more tokens than the transcript for the same second of speech. That is a real cost, paid in context window and in compute, and it buys something specific: the trunk can hear prosody, the rising and falling of pitch and energy that carries emphasis and emotion and that a transcript flattens away, along with background sound and speaker identity, because nothing discarded them on the way in. The suitcase is fuller, but the sound is in it.
Tokens per second of speech for the three front ends. Codec rates come from GenMedia.dsp.tokenRate and the bitrate from GenMedia.dsp.bitrateKbps.
Every residual stage is a token, which is what makes the codebook stack unusual among quality dials: it is also a direct multiplier on the context bill. Switch one more codebook on and the reconstruction gets closer to the original, but the sequence the trunk has to carry grows by one token per frame for every second of audio, so the same knob that sharpens the sound lengthens the prompt. The chart below isolates the quantiser on the same synthetic signal the spectrogram demo uses — a chirp plus a tone at 8 kHz — so the ladder of bars is the error curve of one component rather than of a whole codec. The first codebook removes most of the error, because it sees the whole signal: it quantises the waveform itself, the largest thing there is to describe, and it takes the biggest single bite out of the residual. Each later stage sees only the leftover from the stages before it, and because each uniform quantiser refits its range to that leftover, the residual shrinks by roughly a constant factor every time — about sixteen decibels per codebook here — so the ladder in the chart is a near-straight staircase rather than a curve that levels off. That is what makes the token cost the deciding question: the tenth decibel costs exactly what the first one did, one more token per frame for every second of audio, while the error it removes is already far below what the sound needs. Stopping is therefore a judgement about the codebook count and the frame rate rather than about where a curve flattens: the residual quantiser is worked out in the audio volume's codec part, where the machinery and the reconstruction are the subject, while here the stages matter only for what they cost — how many tokens per second of speech they add, and how much of the context window they consume before the trunk has read a single word.
Residual error after each stage of GenMedia.dsp.rvq, measured in dB below the input variance, so a taller bar means more of the error has been removed. The dashed line is 20 dB below input, and the first stage past it is the one the readout names.
Levels per codebook are fixed at 8, and each stage is one more codebook per frame.
Paralinguistics
What the words do not carry
Paralinguistics is the name for everything about speech that is not the words: pitch and its movement, loudness, speaking rate, pauses, laughter, sighs, and the voice characteristics that identify a speaker. The pitch-and-energy part of that channel has its own name, prosody, the melody and rhythm of an utterance, and a great deal of meaning lives there: the same words can be a question, a warning or a joke depending on how they are delivered. A recognition model is built to be invariant to all of it, so a cascade — recogniser first, trunk second — literally cannot represent it, because the channel was deleted before the trunk ever ran. A native audio model can, because the prosodic information is still in the tokens it was given — but "can" is doing real work in that sentence. Whether the model actually uses that channel depends on the data it was trained on, and on whether anyone asked it to. An architecture capable of hearing tone will still ignore it if every training pair rewarded matching words alone.
Speech instruction following is the task that tests all of this: the model is given a spoken instruction and must answer in speech, and the instruction is often about the delivery rather than the content. Give the model a spoken question and ask it to answer in a particular tone, ask it to describe the emotion in a recording, ask it to whisper back, and a transcript-only model fails at the first hurdle, because the thing it is being asked about — the tone, the whisper — was exactly what the recogniser threw away, so the instruction never reaches it at all. This is also where the boundary blurs with text-to-speech, or TTS, the reverse task of turning text into audio: an answer in speech needs a voice on the way out, and the same token vocabulary can serve both directions. The contours below are a schematic of the channel: two prosodic tracks over one utterance, pitch and energy, flattening as the expressiveness slider goes down, until what is left is barely more than a transcript in the shape of a curve.
Two prosodic contours over one utterance — pitch and energy — from a seeded stream, with the expressiveness slider flattening both. A transcript keeps the words and discards every point on these lines.
Where this shows up
Recognition is one use of a listening model
Recognisers trained on audio alone
The pure recognition task is the domain of models trained end to end on audio and transcripts, with no language model in the loop: frame classifiers with a CTC objective — connectionist temporal classification, a loss that lets a model emit a label per frame without being told in advance where each word starts — transducers, which pair a small predictor with the acoustic model, and encoder-decoder recognisers such as Whisper. That treatment, and the streaming and duplex machinery built on top of it, lives in the audio volume. Two of those words are worth glossing here: streaming means recognising as the audio arrives rather than after it has all been captured, and duplex means listening and speaking at the same time.
Spoken questions, spoken answers
A multimodal trunk that accepts audio tokens can be asked questions about a recording, told to summarise a meeting, or handed a spoken instruction about the sound in front of it. The same model handles both kinds of request, a question about what was said and a question about how it sounded, and they arrive through the same channel and compete for the same context budget. The transcript interface is a special case of this — the case where the audio has already been reduced to words — and a lossy one, which is why the interesting systems keep the audio tokens rather than only their words.
Once a model can hear and speak in the same token space, the remaining problem is timing. Up to this point every task has been a one-shot exchange: all the input arrives, then the model produces all the output. A conversation is not shaped like that. Two streams have to be produced and consumed at once, so the model must keep listening while it is still talking; the user has to be interruptible, which means the model must notice a new voice and abandon a half-finished sentence; and the reply has to start inside a few hundred milliseconds, because a longer pause stops feeling like a conversation. Picture passing notes in a relay, each stage handing the next a short summary rather than the whole pile. That is the real-time loop, and it is the last part of this volume.
Further reading
These papers cover the two front ends, roughly in the order this page introduced them. The first two are self-supervised encoders that learned speech representations before recognition, meaning they were trained on raw audio with no transcripts at all. The third is the recogniser that made attention the default for transcription, the model whose output is the encoder route's transcript. The last two are the audio token models that let a language model hear the signal itself, the ancestry of the native-token route.
- Alexei Baevski, Henry Zhou, Abdelrahman Mohamed and Michael Auli, "wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations", 2020 — pretraining an encoder on raw audio before any transcript exists.
- Wei-Ning Hsu, Benjamin Bolte, Yao-Hung Hubert Tsai and coauthors, "HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units", 2021 — discrete speech units learned without text, the ancestor of codec tokens.
- Alec Radford, Jong Wook Kim, Tao Xu and coauthors, "Robust Speech Recognition via Large-Scale Weak Supervision", 2022 — Whisper, the encoder-decoder recogniser whose output is the encoder route's transcript.
- Zalán Borsos, Raphaël Marinier, Damien Vincent and coauthors, "AudioLM: a Language Modeling Approach to Audio Generation", 2022 — audio and speech as token streams a language model can predict.
- Yunfei Chu, Jin Xu, Xiaohuan Zhou and coauthors, "Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models", 2023 — an audio-language trunk and the instruction-following tasks a transcript interface cannot express.
Cheat sheet
| Term | Meaning here |
|---|---|
| Mel spectrogram | STFT magnitude on a mel frequency axis; the standard audio token grid |
| Window and hop | Analysis trade of frequency against time resolution; the hop sets frames per second |
| Encoder route | A recogniser transcribes first; the trunk receives subwords of text |
| Native-token route | A neural codec emits discrete audio tokens that go into the trunk itself |
| Token rate | frame rate × codebooks; a codec choice, not a property of sound |
| Bitrate | frame rate × codebooks × bits; 75 Hz × 8 × 10 is about 6 kbps |
| Paralinguistics | Pitch, loudness, tempo, pauses, emotion, speaker: everything a transcript discards |
| Speech instruction following | Answering a spoken instruction about a recording, often about its delivery |