Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Pruning and Sparsity

Should knowUnderstand9 minDifficulty

Pruning removes weights (or whole neurons, heads or layers) that contribute little, leaving a sparse network that needs less storage and, if the hardware can exploit the pattern, less computation.

The problem

Trained networks have many near-zero or redundant weights that still cost memory and arithmetic.

The solution

Score weights by importance (often magnitude, or the error their removal causes), set the least important to zero, and retrain or adjust the rest to recover accuracy.

The consequence

Networks can often lose most of their weights with little accuracy loss, but turning zeros into speed needs structure the hardware supports; scattered zeros mostly save storage.

Most weights can go

Han and colleagues trained a network, removed its unimportant connections and retrained the rest: AlexNet went from 61 million to 6.7 million parameters, and VGG-16 from 138 million to 10.3 million, without losing accuracy on ImageNet Established.

Why does this work? The lottery ticket hypothesis proposes that dense networks contain small subnetworks that, trained alone from their original initial weights, reach comparable accuracy Interpretation. Frankle and Carbin found such "winning tickets" at less than 10–20% of the original size for fully connected and convolutional networks on MNIST and CIFAR-10 Established. Whether the picture carries over cleanly to LLMs is a separate question.

At LLM scale

Retraining a 175B-parameter model after pruning is impractical. SparseGPT prunes GPT-family models to at least 50% sparsity in one shot without retraining; on OPT-175B and BLOOM-176B it ran in under 4.5 hours and reached 60% unstructured sparsity with a negligible increase in perplexity Established.

Tiny example. A row of weights [0.8,−0.05,0.02,−0.6][0.8, -0.05, 0.02, -0.6] pruned to 50% by magnitude keeps [0.8,0,0,−0.6][0.8, 0, 0, -0.6]. For an input x=[1,1,1,1]x = [1, 1, 1, 1] the output changes from 0.17 to 0.20: the two dropped weights contributed little. Methods like SparseGPT go further and adjust the kept weights to make up for the dropped ones.

Zeros aren't automatically faster

A GPU multiplies dense blocks. Zeros scattered at random still occupy their slots, so unstructured sparsity mainly saves storage (with a sparse format) rather than time. Structured sparsity removes whole rows, heads or layers, or follows a pattern the hardware understands. SparseGPT also imposes semi-structured patterns such as 2:4, two zeros in every four consecutive weights, a format its authors note delivers speed-ups on NVIDIA's Ampere GPUs, at some accuracy cost for smaller models Established.

What to remember

  • Magnitude pruning: drop the smallest weights, retrain the rest.
  • AlexNet: 61M → 6.7M parameters (9×) without accuracy loss.
  • Lottery tickets: subnetworks 10–20% of the size train as well from their original initialization (small vision networks).
  • SparseGPT: 50–60% of a 175B model's weights removed in one shot, no retraining.
  • Unstructured zeros rarely speed up GPUs; structured patterns (e.g. 2 of every 4) can.

Key papers

Optional

The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Jonathan Frankle, Michael Carbin · 2018

A provocative finding about why pruning works: dense networks contain small subnetworks that could have been trained alone.

How to read it: Read the introduction and Section 2. The experiments are on small vision networks; whether it carries over to LLMs is a separate question.

~40 min readarXiv:1803.03635✓ verified 2026-10-05

Watch