Concept · Chapter 11: Inside Modern LLMs
Pruning and Sparsity
Pruning removes weights (or whole neurons, heads or layers) that contribute little, leaving a sparse network that needs less storage and, if the hardware can exploit the pattern, less computation.
The problem
Trained networks have many near-zero or redundant weights that still cost memory and arithmetic.
The solution
Score weights by importance (often magnitude, or the error their removal causes), set the least important to zero, and retrain or adjust the rest to recover accuracy.
The consequence
Networks can often lose most of their weights with little accuracy loss, but turning zeros into speed needs structure the hardware supports; scattered zeros mostly save storage.
You should understand first
- Vectors
- Dot Product
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Loss Functions
- Derivatives and Gradients
- Gradient Descent
- Linear Regression
- Probability and Distributions
- Entropy
- Softmax
- Cross-Entropy Loss
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Matrix Multiplication
- Multilayer Perceptron (MLP)
- Expected Value and Variance
- Sampling and Uncertainty
- Generalization, Overfitting and Underfitting
- Regularization
- Pruning and Sparsity
Most weights can go
Han and colleagues trained a network, removed its unimportant connections and retrained the rest: AlexNet went from 61 million to 6.7 million parameters, and VGG-16 from 138 million to 10.3 million, without losing accuracy on ImageNet Established.
Why does this work? The lottery ticket hypothesis proposes that dense networks contain small subnetworks that, trained alone from their original initial weights, reach comparable accuracy Interpretation. Frankle and Carbin found such "winning tickets" at less than 10–20% of the original size for fully connected and convolutional networks on MNIST and CIFAR-10 Established. Whether the picture carries over cleanly to LLMs is a separate question.
At LLM scale
Retraining a 175B-parameter model after pruning is impractical. SparseGPT prunes GPT-family models to at least 50% sparsity in one shot without retraining; on OPT-175B and BLOOM-176B it ran in under 4.5 hours and reached 60% unstructured sparsity with a negligible increase in perplexity Established.
Tiny example. A row of weights pruned to 50% by magnitude keeps . For an input the output changes from 0.17 to 0.20: the two dropped weights contributed little. Methods like SparseGPT go further and adjust the kept weights to make up for the dropped ones.
Zeros aren't automatically faster
A GPU multiplies dense blocks. Zeros scattered at random still occupy their slots, so unstructured sparsity mainly saves storage (with a sparse format) rather than time. Structured sparsity removes whole rows, heads or layers, or follows a pattern the hardware understands. SparseGPT also imposes semi-structured patterns such as 2:4, two zeros in every four consecutive weights, a format its authors note delivers speed-ups on NVIDIA's Ampere GPUs, at some accuracy cost for smaller models Established.
What to remember
- Magnitude pruning: drop the smallest weights, retrain the rest.
- AlexNet: 61M → 6.7M parameters (9×) without accuracy loss.
- Lottery tickets: subnetworks 10–20% of the size train as well from their original initialization (small vision networks).
- SparseGPT: 50–60% of a 175B model's weights removed in one shot, no retraining.
- Unstructured zeros rarely speed up GPUs; structured patterns (e.g. 2 of every 4) can.
Key papers
Learning both Weights and Connections for Efficient Neural Networks
Song Han, Jeff Pool et al. · 2015
The classic train–prune–retrain recipe that showed most connections in a trained network can be removed.
The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks
Jonathan Frankle, Michael Carbin · 2018
A provocative finding about why pruning works: dense networks contain small subnetworks that could have been trained alone.
How to read it: Read the introduction and Section 2. The experiments are on small vision networks; whether it carries over to LLMs is a separate question.
SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
Elias Frantar, Dan Alistarh · 2023
Showed that GPT-scale models can lose half their weights in one pass, without retraining.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.