Concept · Chapter 15: Multimodal AI
Audio Language Models and Voice
Audio language models treat sound as a token sequence, using spectrogram features to listen and neural-codec codes to speak, so one model can hear speech, music and tone and answer in audio without converting everything to text.
The problem
A voice assistant built from speech-to-text, a text LLM and text-to-speech loses tone, emotion, speakers and background sound at the first step, and the three stages add up to seconds of delay.
The solution
Turn audio into tokens a Transformer can read (encoder features from spectrograms) and write (discrete codes from a neural audio codec), and train one model across speech and text.
The consequence
Conversational latency approaches human turn-taking and the model can hear and produce non-verbal sound, but voice brings new risks such as voice imitation and speaker identification.
You should understand first
- Vectors
- Audio and Spectrograms
- Speech Recognition and Synthesis
- Text as Data
- One-Hot Encoding
- Tokenization
- Tensors and Shapes
- Images as Tensors
- Dot Product
- Convolution
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Chain Rule
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Convolutional Neural Networks
- Embeddings
- Attention
- Self-Attention
- Vision Transformer (ViT)
- Turning Signals into Tokens
- Audio Language Models and Voice
Hearing: from waveform to tokens
Chapter 5 turned sound into a spectrogram, a picture of frequency over time. Speech recognition then mapped it to text. Whisper splits audio into 30-second segments, computes an 80-channel log-magnitude mel spectrogram on 25 ms windows with a 10 ms stride, and passes it through a small convolutional stem whose second layer has stride 2 before the Transformer encoder. Established Arithmetic from those settings: 30 seconds is 3,000 frames, halved to 1,500 encoder positions, about 50 per second.
Speaking: neural audio codecs
To generate audio token by token, a model needs discrete audio tokens that can be turned back into sound. SoundStream is a neural audio codec: a convolutional encoder and decoder around a residual vector quantizer, trained end to end; one model covers 3 to 18 kbps, and at 3 kbps it outperformed Opus at 12 kbps in listening tests. Established
Residual vector quantization is the key trick: quantize a frame with one codebook, subtract the chosen code, quantize what is left with a second codebook, and so on. Each extra codebook adds detail and bitrate. It is the audio version of the image codebook in the tokens lab, stacked.
AudioLM treated audio generation as language modelling over a hybrid of tokens: discretized features of a self-supervised speech model for long-term structure and neural-codec codes for fidelity; trained on speech without transcripts, it continued speech while keeping the speaker's identity. EstablishedOne model instead of three
OpenAI's GPT-4o announcement (May 2024) says its earlier Voice Mode chained three models (transcription, GPT-3.5 or GPT-4, text-to-speech) with average latencies of 2.8 and 5.4 seconds, and that this pipeline could not observe tone, multiple speakers or background noise, or output laughter or singing. Established GPT-4o was trained end to end across text, vision and audio, and responds to audio in as little as 232 ms, with an average of 320 ms. Established These are the company's own measurements.
The general lesson is the one this chapter keeps returning to: converting a modality to text early is simple and inspectable, but it throws away whatever text cannot express. InterpretationNew risks
The GPT-4o system card focuses its safety evaluation on speech-to-speech, including risks such as unauthorized voice generation and speaker identification. Established Voice is biometric: a model that can imitate any voice, or recognize who is speaking, needs constraints that a text model never did.
Tiny example
At about 50 encoder positions per second, a 10-minute meeting recording is roughly 30,000 positions before any text: more than many models' whole context windows a few years ago. Summarizing long audio therefore usually means chunking, the same way RAG chunks long documents.
Mini experiment
A codec with 8 residual codebooks of 1,024 entries each produces one code per codebook per frame, at 75 frames per second. How many audio tokens per second does a model that predicts every code have to generate, and how many bits per second does that represent? Compare with ordinary speech, roughly 150 words a minute.
What to remember
- Listening: spectrogram frames through an encoder (Whisper: 30-second chunks, 80-channel log-mel, 10 ms stride).
- Speaking: a neural codec turns audio into discrete codes and back; residual vector quantization stacks codebooks for detail.
- SoundStream at 3 kbps beat the Opus codec at 12 kbps in listening tests.
- AudioLM combined coarse 'semantic' tokens for structure with codec tokens for fidelity.
- GPT-4o replaced a three-model voice pipeline (2.8–5.4 s average latency) with one network (about 0.32 s average).
Key papers
Robust Speech Recognition via Large-Scale Weak Supervision
Alec Radford, Jong Wook Kim et al. · 2022
Whisper: an encoder–decoder Transformer trained on 680,000 hours of audio paired with transcripts gathered from the internet. Speech recognition became one more sequence-to-sequence problem solved by scale.
SoundStream: An End-to-End Neural Audio Codec
Neil Zeghidour, Alejandro Luebs et al. · 2021
A neural codec that turns audio into a short stream of discrete codes. Codecs like it are how audio language models read and write sound as tokens.
How to read it: Residual vector quantization is the idea to take away: quantize, subtract, quantize the remainder again. Each extra codebook adds detail and bitrate.
AudioLM: a Language Modeling Approach to Audio Generation
Zalán Borsos, Raphaël Marinier et al. · 2022
Treated audio generation as language modelling over audio tokens, continuing speech and piano without transcripts or scores.
How to read it: The comparison of coarse 'semantic' tokens and fine 'acoustic' tokens is the useful idea; it reappears in later speech models.
High Fidelity Neural Audio Compression
Alexandre Défossez, Jade Copet et al. · 2022
EnCodec, an open neural audio codec whose discrete codes are widely used as audio tokens in research.
How to read it: Skim for the bitrate and code-rate figures: they tell you how many audio tokens per second a model built on these codes must handle.
GPT-4o System Card
OpenAI et al. · 2024
Documents a model trained end to end across text, vision and audio, with a focus on the new risks of speech-to-speech interaction.
How to read it: Read the speech-specific risk sections (voice imitation, speaker identification): new modalities bring new failure modes, not just new features.