Concept · Chapter 11: Inside Modern LLMs
Quantization
Quantization stores weights (and sometimes activations or the KV cache) in fewer bits, such as 8 or 4 instead of 16, by mapping each group of numbers onto a small grid of levels with a shared scale, trading a little accuracy for much less memory and memory traffic.
The problem
A 70B model needs 140 GB just for its 16-bit weights, more than one 80 GB GPU, and every decode step must read all of it.
The solution
Replace each weight by a small integer times a scale shared by a group of weights. Pick the scale per small group so one large value can't ruin the rest, keep troublesome outliers in higher precision, and correct rounding errors with a little calibration data.
The consequence
8-bit weights are close to lossless for many models and 4-bit weights are widely used for inference, letting large models run on fewer or smaller GPUs and decode faster. Below about 4 bits, quality loss becomes harder to avoid.
You should understand first
- Derivatives and Gradients
- Loss Functions
- Gradient Descent
- Probability and Distributions
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
- Mixed-Precision Training (FP16 and BF16)
- Text as Data
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Embeddings
- Attention
- Self-Attention
- Causal Masking
- Autoregressive Next-Token Prediction
- Multi-Head Attention
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Prefill, Decode and the Memory Wall
- Quantization
A coarser ruler
A 16-bit number can take about 65,000 values. A 4-bit integer can take 16. Quantization picks a grid of levels for a group of weights and rounds each weight to the nearest level.
The simplest scheme, absmax, makes the grid symmetric around zero and just wide enough for the largest weight in the group:
Tiny example. Four weights in 4 bits: the integer levels run from −7 to 7, so . Dividing and rounding gives , which decode to . Each weight now needs 4 bits plus a share of the scale.
Why one large number is a problem
The scale is set by the largest value. If one weight in a group is 40 times larger than the others, the grid stretches to fit it, and almost every ordinary weight lies closer to zero than to the first level: it rounds to zero.
Dettmers and colleagues found that large language models develop rare but systematic outlier features with very large magnitudes, which ruin quantization precision once they appear in all Transformer layers at about 6.7B parameters Established. LLM.int8() uses separate scaling constants for each inner product in a matrix multiplication and keeps the outlier dimensions in a 16-bit multiplication while more than 99.9% of values are multiplied in 8 bits; with it, models of up to 175B parameters ran in 8 bits without performance degradation Established.
The general fix is smaller groups: one scale per block of, say, 64 weights confines an outlier's damage to its own block. Scales cost memory: with 32-bit scales and blocks of 64, the scales add 32/64 = 0.5 bits per parameter, which QLoRA's double quantization reduces to about 0.127 Established.
Try it · toy model
Round a row of weights to 8, 4, 3 or 2 bits, add one outlier, and see why a shared scale fails and per-group scales rescue it.
Rounding smarter
Rounding each weight to its nearest level independently ignores that errors add up in a layer's output.
- GPTQ quantizes weights a block of columns at a time and updates the remaining weights to compensate for the rounding error, using approximate second-order information; it quantized 175B-parameter models to 3–4 bits in about four GPU hours with negligible accuracy loss Established.
- AWQ observes that protecting about 1% of salient weights, identified from the activations rather than the weights, greatly reduces quantization error, and protects them by scaling those channels up before quantizing Established.
Floating point with fewer bits
Integers are not the only option. FP8 comes in two encodings, E4M3 (4 exponent bits, 3 mantissa bits) and E5M2 (5 and 2), and matched 16-bit training quality on the models tested, up to 175B parameters Established. Meta served Llama 3 405B with FP8 quantization of most parameters and activations in the feed-forward layers, leaving the self-attention layers unquantized, and reported up to 50% higher throughput in the prefill stage Established.
What it buys, and what it doesn't
A decode step reads every weight (prefill and decode), so halving the bytes per weight roughly halves that part of the step, if the hardware has fast kernels for the format. Memory falls in step: Llama 3 70B needs 140 GB in 16 bits, 70 GB in 8 bits and about 39 GB at 4 bits with 64-weight groups.
Why should I care?
As a researcher
How few bits a network needs, and why a few outlier features break naive schemes, says something about how large models represent information.
As an engineer
Quantization is usually the first lever for fitting a model on your hardware: halving the bytes per weight roughly halves memory and the decode step's weight-reading time.
Modern systems that depend on it
- QLoRA
- running LLMs locally
- FP8 serving
- KV-cache quantization
Historical context
Before
Models were served in 32- or 16-bit floating point; lower precision was associated with small vision models and specialized hardware.
After
8-bit and 4-bit formats are routine for LLM inference; 16-bit is the reference and FP8 is used on hardware that supports it.
Used today
Open models are commonly distributed and run in 8-bit and 4-bit formats, and Meta reported FP8 inference for Llama 3 405B.
What to remember
- Absmax quantization: scale = max|w| ÷ (2^(bits−1) − 1), store round(w ÷ scale).
- One outlier inflates a shared scale and flattens everything else; per-group scales (e.g. 64 weights) contain the damage.
- Each 32-bit scale per 64 weights costs 0.5 extra bits per weight.
- LLM.int8(): outlier features from ~6.7B parameters; keep them in 16-bit, the other 99.9% in 8-bit.
- GPTQ: 175B to 3–4 bits in about 4 GPU hours, by compensating rounding errors column by column.
Key papers
DeepSeek-V3 Technical Report
DeepSeek-AI et al. · 2024
A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.
How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.
LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale
Tim Dettmers, Mike Lewis et al. · 2022
Found why naive 8-bit quantization breaks large models (a few huge outlier features) and how to work around them, halving inference memory without degradation.
How to read it: Figure 1 (accuracy collapse at 6.7B) and Section 4 on emergent outlier features are the memorable parts.
GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Elias Frantar, Saleh Ashkboos et al. · 2022
Made 3–4-bit weights practical for very large models without retraining, which is how many open models are run on single GPUs.
How to read it: Section 4 builds the algorithm step by step; Figure 2 and Algorithm 1 show the block-by-block procedure.
AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration
Ji Lin, Jiaming Tang et al. · 2023
Shows that which weights matter depends on the activations they meet, and protects those with a simple rescaling.
How to read it: Section 3.1 (protecting 1% of weights) and 3.2 (the scaling trick) are the method.
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni et al. · 2023
Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.
How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.
FP8 Formats for Deep Learning
Paulius Micikevicius, Dusan Stosic et al. · 2022
Specified the two 8-bit floating-point formats (E4M3 and E5M2) that recent GPUs support for training and inference.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.