Concept · Chapter 11: Inside Modern LLMs
LoRA, QLoRA and Parameter-Efficient Fine-Tuning
LoRA fine-tunes a large model by freezing its weights and learning, for chosen weight matrices, a small low-rank update (the product of two thin matrices), so only a tiny fraction of parameters is trained and stored per task.
The problem
Full fine-tuning updates every parameter, needs optimizer states for all of them, and produces a full-size copy of the model for every task.
The solution
Write each adapted weight as W₀ + BA, with W₀ frozen, B (d × r) and A (r × k) trainable and r much smaller than d. Train only A and B; merge BA into W₀ for serving, or swap adapters per task. QLoRA also stores W₀ in 4 bits.
The consequence
Fine-tuning became possible on a single GPU and adapters became cheap to store and swap. The update is constrained to a low rank, which is usually enough for adapting style and tasks but not necessarily for teaching a model large amounts of new knowledge.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Multi-Head Attention
- Causal Masking
- Positional Encoding
- Residual Connections
- Layer Normalization
- Feed-Forward Sublayer (MLP)
- The Transformer Block
- Encoder, Decoder & Encoder–Decoder
- Pretrain, Then Fine-Tune
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Supervised Fine-Tuning
- Matrix Multiplication
- Derivatives and Gradients
- Gradient Descent
- Expected Value and Variance
- Stochastic Gradient Descent (SGD)
- The Chain Rule
- The Turing Test
- Symbolic AI
- Logic and Rules
- Expert Systems
- Knowledge Representation
- The Knowledge-Acquisition Bottleneck
- From Rules to Learning
- Supervised, Unsupervised and Self-Supervised Learning
- Features, Labels and Tasks
- Linear Regression
- Logistic Regression
- The Perceptron
- Activation Functions
- The Artificial Neuron
- Multilayer Perceptron (MLP)
- The Forward Pass
- Computational Graphs and Autodiff
- Backpropagation
- Momentum and Adam
- The Pretraining Loop
- Mixed-Precision Training (FP16 and BF16)
- One-Hot Encoding
- Tokenization
- Autoregressive Next-Token Prediction
- Pretraining at Scale
- Parameters, Tokens and Context Windows
- Where Training Memory Goes
- LoRA, QLoRA and Parameter-Efficient Fine-Tuning
Fine-tuning a giant is mostly overhead
Supervised fine-tuning changes every weight. For a model with parameters, training needs about 16 bytes per parameter for weights, gradients and Adam's states (training memory), and the result is a new -parameter model. For GPT-3 175B, the LoRA authors note, each fine-tuned instance would be a full 175B-parameter model, prohibitively expensive to deploy for many tasks Established.
Earlier, Houlsby and colleagues inserted small trainable adapter layers into a frozen BERT and came within 0.4% of full fine-tuning on GLUE while adding 3.6% parameters per task Established. But the extra layers add a little computation at every step of inference.
A low-rank update
LoRA freezes the pretrained weights and injects trainable rank-decomposition matrices: the update to a weight matrix is , with initialized randomly and at zero, so the model starts out unchanged, and the update is scaled by Established:
Tiny example. A 4,096 × 4,096 projection has 16.8 million weights. A rank-8 update trains numbers: 0.39% of the matrix. Adam's states are needed only for those.
- This matrix, full fine-tuning
- 16,777,216
- This matrix, LoRA
- 65,536 (0.39%)
- Llama 3 8B, query + value in all 32 layers
- 3.4M (0.043% of 8B)
Trainable numbers are r × (d + k) per matrix. B starts at zero, so training begins from the unchanged model. Llama 3 8B’s value projection is 4,096 × 1,024 because of grouped-query attention. Counts are our arithmetic from the published architecture; Adam’s states are needed only for the trainable part.
Compared with fine-tuning GPT-3 175B with Adam, LoRA reduced the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times, while matching or beating fine-tuning quality on the models tested Established. With rank 4 applied only to the query and value projections, the checkpoint shrank from 350 GB to about 35 MB; storing 100 adapted models takes about 354 GB instead of 35 TB Established.
No extra latency, or many tasks at once
For serving, can be computed and stored, so inference runs as usual with no added latency; switching tasks means subtracting one and adding another Established. Alternatively, keep the adapters separate and let one copy of the base model serve requests for many adapters.
QLoRA: a 4-bit base
LoRA still keeps the frozen base in 16 bits. QLoRA backpropagates through a frozen 4-bit base model into LoRA adapters, which reduced memory enough to fine-tune a 65B-parameter model on a single 48 GB GPU while preserving 16-bit fine-tuning task performance Established. It introduced the 4-bit NormalFloat (NF4) data type for normally distributed weights, double quantization of the quantization constants, and paged optimizers to absorb memory spikes Established. See quantization for why the constants matter.
Why should I care?
As a researcher
That a rank-4 or rank-8 update suffices for many adaptations says the useful change from fine-tuning is low-dimensional, an observation the LoRA paper investigates.
As an engineer
LoRA and QLoRA are how most teams fine-tune open models on modest hardware, and how one base model can serve many customers with small per-customer adapters.
Modern systems that depend on it
- QLoRA
- adapter serving
- instruction tuning on one GPU
- domain adaptation
Historical context
Before
Adapting a pretrained model meant fine-tuning all its weights and storing a full copy per task, or adding adapter layers that slowed inference.
After
Frozen base model plus small trainable low-rank updates; with QLoRA, the frozen base is stored in 4 bits.
Used today
Parameter-efficient fine-tuning, mostly LoRA and its variants, is a standard option in fine-tuning libraries and services for open models.
What to remember
- W = W₀ + (α/r)·BA; W₀ frozen, B is d × r, A is r × k.
- B starts at zero, so training starts from the original model.
- Trainable parameters per matrix: r(d + k) instead of dk (0.4% for a 4,096 × 4,096 matrix at r = 8).
- Merge BA into W₀ for serving: no extra latency. Keep it separate to swap tasks.
- QLoRA: 4-bit NF4 base + LoRA; a 65B model fine-tuned on one 48 GB GPU.
Key papers
QLoRA: Efficient Finetuning of Quantized LLMs
Tim Dettmers, Artidoro Pagnoni et al. · 2023
Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.
How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen et al. · 2021
The standard way to adapt a large model cheaply: freeze its weights and learn a small low-rank update for some matrices.
How to read it: Section 4 (the method, one page) and Section 7 (why a low rank is enough) are the essentials.
Parameter-Efficient Transfer Learning for NLP
Neil Houlsby, Andrei Giurgiu et al. · 2019
Introduced adapters, small trainable layers inserted into a frozen network: the start of parameter-efficient fine-tuning.
How to read it: Figure 2 shows where the adapters go; that picture is most of the paper.
Watch
Stanford Online
Stanford CS336 Language Modeling from Scratch | Spring 2025 | Lecture 10: Inference
The closest single lecture to this chapter: the arithmetic of serving a language model and the main ways to make it cheaper.