Skip to content
Road to Intelligence

Concept · Chapter 15: Multimodal AI

Audio Language Models and Voice

Should knowUnderstand11 minDifficulty

Audio language models treat sound as a token sequence, using spectrogram features to listen and neural-codec codes to speak, so one model can hear speech, music and tone and answer in audio without converting everything to text.

The problem

A voice assistant built from speech-to-text, a text LLM and text-to-speech loses tone, emotion, speakers and background sound at the first step, and the three stages add up to seconds of delay.

The solution

Turn audio into tokens a Transformer can read (encoder features from spectrograms) and write (discrete codes from a neural audio codec), and train one model across speech and text.

The consequence

Conversational latency approaches human turn-taking and the model can hear and produce non-verbal sound, but voice brings new risks such as voice imitation and speaker identification.

Hearing: from waveform to tokens

Chapter 5 turned sound into a spectrogram, a picture of frequency over time. Speech recognition then mapped it to text. Whisper splits audio into 30-second segments, computes an 80-channel log-magnitude mel spectrogram on 25 ms windows with a 10 ms stride, and passes it through a small convolutional stem whose second layer has stride 2 before the Transformer encoder. Established Arithmetic from those settings: 30 seconds is 3,000 frames, halved to 1,500 encoder positions, about 50 per second.

Speaking: neural audio codecs

To generate audio token by token, a model needs discrete audio tokens that can be turned back into sound. SoundStream is a neural audio codec: a convolutional encoder and decoder around a residual vector quantizer, trained end to end; one model covers 3 to 18 kbps, and at 3 kbps it outperformed Opus at 12 kbps in listening tests. Established

Residual vector quantization is the key trick: quantize a frame with one codebook, subtract the chosen code, quantize what is left with a second codebook, and so on. Each extra codebook adds detail and bitrate. It is the audio version of the image codebook in the tokens lab, stacked.

AudioLM treated audio generation as language modelling over a hybrid of tokens: discretized features of a self-supervised speech model for long-term structure and neural-codec codes for fidelity; trained on speech without transcripts, it continued speech while keeping the speaker's identity. Established

One model instead of three

OpenAI's GPT-4o announcement (May 2024) says its earlier Voice Mode chained three models (transcription, GPT-3.5 or GPT-4, text-to-speech) with average latencies of 2.8 and 5.4 seconds, and that this pipeline could not observe tone, multiple speakers or background noise, or output laughter or singing. Established GPT-4o was trained end to end across text, vision and audio, and responds to audio in as little as 232 ms, with an average of 320 ms. Established These are the company's own measurements.

The general lesson is the one this chapter keeps returning to: converting a modality to text early is simple and inspectable, but it throws away whatever text cannot express. Interpretation

New risks

The GPT-4o system card focuses its safety evaluation on speech-to-speech, including risks such as unauthorized voice generation and speaker identification. Established Voice is biometric: a model that can imitate any voice, or recognize who is speaking, needs constraints that a text model never did.

Tiny example

At about 50 encoder positions per second, a 10-minute meeting recording is roughly 30,000 positions before any text: more than many models' whole context windows a few years ago. Summarizing long audio therefore usually means chunking, the same way RAG chunks long documents.

Mini experiment

A codec with 8 residual codebooks of 1,024 entries each produces one code per codebook per frame, at 75 frames per second. How many audio tokens per second does a model that predicts every code have to generate, and how many bits per second does that represent? Compare with ordinary speech, roughly 150 words a minute.

What to remember

  • Listening: spectrogram frames through an encoder (Whisper: 30-second chunks, 80-channel log-mel, 10 ms stride).
  • Speaking: a neural codec turns audio into discrete codes and back; residual vector quantization stacks codebooks for detail.
  • SoundStream at 3 kbps beat the Opus codec at 12 kbps in listening tests.
  • AudioLM combined coarse 'semantic' tokens for structure with codec tokens for fidelity.
  • GPT-4o replaced a three-model voice pipeline (2.8–5.4 s average latency) with one network (about 0.32 s average).

Key papers

Important

Robust Speech Recognition via Large-Scale Weak Supervision

Alec Radford, Jong Wook Kim et al. · 2022

Whisper: an encoder–decoder Transformer trained on 680,000 hours of audio paired with transcripts gathered from the internet. Speech recognition became one more sequence-to-sequence problem solved by scale.

~40 min readarXiv:2212.04356✓ verified 2026-09-26
Optional

SoundStream: An End-to-End Neural Audio Codec

Neil Zeghidour, Alejandro Luebs et al. · 2021

A neural codec that turns audio into a short stream of discrete codes. Codecs like it are how audio language models read and write sound as tokens.

How to read it: Residual vector quantization is the idea to take away: quantize, subtract, quantize the remainder again. Each extra codebook adds detail and bitrate.

~35 min readarXiv:2107.03312✓ verified 2026-10-07
Optional

AudioLM: a Language Modeling Approach to Audio Generation

Zalán Borsos, Raphaël Marinier et al. · 2022

Treated audio generation as language modelling over audio tokens, continuing speech and piano without transcripts or scores.

How to read it: The comparison of coarse 'semantic' tokens and fine 'acoustic' tokens is the useful idea; it reappears in later speech models.

~35 min readarXiv:2209.03143✓ verified 2026-10-07
Optional

High Fidelity Neural Audio Compression

Alexandre Défossez, Jade Copet et al. · 2022

EnCodec, an open neural audio codec whose discrete codes are widely used as audio tokens in research.

How to read it: Skim for the bitrate and code-rate figures: they tell you how many audio tokens per second a model built on these codes must handle.

~35 min readarXiv:2210.13438✓ verified 2026-10-07
Optional

GPT-4o System Card

OpenAI et al. · 2024

Documents a model trained end to end across text, vision and audio, with a focus on the new risks of speech-to-speech interaction.

How to read it: Read the speech-specific risk sections (voice imitation, speaker identification): new modalities bring new failure modes, not just new features.

~45 min readarXiv:2410.21276✓ verified 2026-10-07