Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Quantization

Must knowKnow well15 minDifficulty

Quantization stores weights (and sometimes activations or the KV cache) in fewer bits, such as 8 or 4 instead of 16, by mapping each group of numbers onto a small grid of levels with a shared scale, trading a little accuracy for much less memory and memory traffic.

The problem

A 70B model needs 140 GB just for its 16-bit weights, more than one 80 GB GPU, and every decode step must read all of it.

The solution

Replace each weight by a small integer times a scale shared by a group of weights. Pick the scale per small group so one large value can't ruin the rest, keep troublesome outliers in higher precision, and correct rounding errors with a little calibration data.

The consequence

8-bit weights are close to lossless for many models and 4-bit weights are widely used for inference, letting large models run on fewer or smaller GPUs and decode faster. Below about 4 bits, quality loss becomes harder to avoid.

A coarser ruler

A 16-bit number can take about 65,000 values. A 4-bit integer can take 16. Quantization picks a grid of levels for a group of weights and rounds each weight to the nearest level.

The simplest scheme, absmax, makes the grid symmetric around zero and just wide enough for the largest weight in the group:

s=max⁡i∣wi∣2b−1−1,qi=round ⁣(wis),w^i=qi s.s = \frac{\max_i |w_i|}{2^{b-1} - 1}, \qquad q_i = \text{round}\!\left(\frac{w_i}{s}\right), \qquad \hat w_i = q_i \, s .

Tiny example. Four weights [0.12,−0.50,0.31,0.04][0.12, -0.50, 0.31, 0.04] in 4 bits: the integer levels run from −7 to 7, so s=0.50/7≈0.0714s = 0.50 / 7 \approx 0.0714. Dividing and rounding gives q=[2,−7,4,1]q = [2, -7, 4, 1], which decode to [0.143,−0.500,0.286,0.071][0.143, -0.500, 0.286, 0.071]. Each weight now needs 4 bits plus a share of the scale.

Why one large number is a problem

The scale is set by the largest value. If one weight in a group is 40 times larger than the others, the grid stretches to fit it, and almost every ordinary weight lies closer to zero than to the first level: it rounds to zero.

Dettmers and colleagues found that large language models develop rare but systematic outlier features with very large magnitudes, which ruin quantization precision once they appear in all Transformer layers at about 6.7B parameters Established. LLM.int8() uses separate scaling constants for each inner product in a matrix multiplication and keeps the outlier dimensions in a 16-bit multiplication while more than 99.9% of values are multiplied in 8 bits; with it, models of up to 175B parameters ran in 8 bits without performance degradation Established.

The general fix is smaller groups: one scale per block of, say, 64 weights confines an outlier's damage to its own block. Scales cost memory: with 32-bit scales and blocks of 64, the scales add 32/64 = 0.5 bits per parameter, which QLoRA's double quantization reduces to about 0.127 Established.

Try it · toy model

Squeeze the Weights

Round a row of weights to 8, 4, 3 or 2 bits, add one outlier, and see why a shared scale fails and per-group scales rescue it.

Know well7 min

Rounding smarter

Rounding each weight to its nearest level independently ignores that errors add up in a layer's output.

  • GPTQ quantizes weights a block of columns at a time and updates the remaining weights to compensate for the rounding error, using approximate second-order information; it quantized 175B-parameter models to 3–4 bits in about four GPU hours with negligible accuracy loss Established.
  • AWQ observes that protecting about 1% of salient weights, identified from the activations rather than the weights, greatly reduces quantization error, and protects them by scaling those channels up before quantizing Established.

Floating point with fewer bits

Integers are not the only option. FP8 comes in two encodings, E4M3 (4 exponent bits, 3 mantissa bits) and E5M2 (5 and 2), and matched 16-bit training quality on the models tested, up to 175B parameters Established. Meta served Llama 3 405B with FP8 quantization of most parameters and activations in the feed-forward layers, leaving the self-attention layers unquantized, and reported up to 50% higher throughput in the prefill stage Established.

What it buys, and what it doesn't

A decode step reads every weight (prefill and decode), so halving the bytes per weight roughly halves that part of the step, if the hardware has fast kernels for the format. Memory falls in step: Llama 3 70B needs 140 GB in 16 bits, 70 GB in 8 bits and about 39 GB at 4 bits with 64-weight groups.

Why should I care?

As a researcher

How few bits a network needs, and why a few outlier features break naive schemes, says something about how large models represent information.

As an engineer

Quantization is usually the first lever for fitting a model on your hardware: halving the bytes per weight roughly halves memory and the decode step's weight-reading time.

Modern systems that depend on it

  • QLoRA
  • running LLMs locally
  • FP8 serving
  • KV-cache quantization

Historical context

Before

Models were served in 32- or 16-bit floating point; lower precision was associated with small vision models and specialized hardware.

After

8-bit and 4-bit formats are routine for LLM inference; 16-bit is the reference and FP8 is used on hardware that supports it.

Used today

Open models are commonly distributed and run in 8-bit and 4-bit formats, and Meta reported FP8 inference for Llama 3 405B.

What to remember

  • Absmax quantization: scale = max|w| ÷ (2^(bits−1) − 1), store round(w ÷ scale).
  • One outlier inflates a shared scale and flattens everything else; per-group scales (e.g. 64 weights) contain the damage.
  • Each 32-bit scale per 64 weights costs 0.5 extra bits per weight.
  • LLM.int8(): outlier features from ~6.7B parameters; keep them in 16-bit, the other 99.9% in 8-bit.
  • GPTQ: 175B to 3–4 bits in about 4 GPU hours, by compensating rounding errors column by column.

Key papers

Important

DeepSeek-V3 Technical Report

DeepSeek-AI et al. · 2024

A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.

How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.

~1 h 30 min readarXiv:2412.19437✓ verified 2026-10-05
Essential

LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale

Tim Dettmers, Mike Lewis et al. · 2022

Found why naive 8-bit quantization breaks large models (a few huge outlier features) and how to work around them, halving inference memory without degradation.

How to read it: Figure 1 (accuracy collapse at 6.7B) and Section 4 on emergent outlier features are the memorable parts.

~50 min readarXiv:2208.07339✓ verified 2026-10-05
Important

GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Elias Frantar, Saleh Ashkboos et al. · 2022

Made 3–4-bit weights practical for very large models without retraining, which is how many open models are run on single GPUs.

How to read it: Section 4 builds the algorithm step by step; Figure 2 and Algorithm 1 show the block-by-block procedure.

~45 min readarXiv:2210.17323✓ verified 2026-10-05
Optional

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration

Ji Lin, Jiaming Tang et al. · 2023

Shows that which weights matter depends on the activations they meet, and protects those with a simple rescaling.

How to read it: Section 3.1 (protecting 1% of weights) and 3.2 (the scaling trick) are the method.

~35 min readarXiv:2306.00978✓ verified 2026-10-05
Essential

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni et al. · 2023

Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.

How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.

~45 min readarXiv:2305.14314✓ verified 2026-10-05
Optional

FP8 Formats for Deep Learning

Paulius Micikevicius, Dusan Stosic et al. · 2022

Specified the two 8-bit floating-point formats (E4M3 and E5M2) that recent GPUs support for training and inference.

~20 min readarXiv:2209.05433✓ verified 2026-10-05

Watch