Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

LoRA, QLoRA and Parameter-Efficient Fine-Tuning

Must knowKnow well14 minDifficulty

LoRA fine-tunes a large model by freezing its weights and learning, for chosen weight matrices, a small low-rank update (the product of two thin matrices), so only a tiny fraction of parameters is trained and stored per task.

The problem

Full fine-tuning updates every parameter, needs optimizer states for all of them, and produces a full-size copy of the model for every task.

The solution

Write each adapted weight as W₀ + BA, with W₀ frozen, B (d × r) and A (r × k) trainable and r much smaller than d. Train only A and B; merge BA into W₀ for serving, or swap adapters per task. QLoRA also stores W₀ in 4 bits.

The consequence

Fine-tuning became possible on a single GPU and adapters became cheap to store and swap. The update is constrained to a low rank, which is usually enough for adapting style and tasks but not necessarily for teaching a model large amounts of new knowledge.

Fine-tuning a giant is mostly overhead

Supervised fine-tuning changes every weight. For a model with NN parameters, training needs about 16 bytes per parameter for weights, gradients and Adam's states (training memory), and the result is a new NN-parameter model. For GPT-3 175B, the LoRA authors note, each fine-tuned instance would be a full 175B-parameter model, prohibitively expensive to deploy for many tasks Established.

Earlier, Houlsby and colleagues inserted small trainable adapter layers into a frozen BERT and came within 0.4% of full fine-tuning on GLUE while adding 3.6% parameters per task Established. But the extra layers add a little computation at every step of inference.

A low-rank update

LoRA freezes the pretrained weights and injects trainable rank-decomposition matrices: the update to a weight matrix is ΔW=BA\Delta W = BA, with AA initialized randomly and BB at zero, so the model starts out unchanged, and the update is scaled by α/r\alpha / r Established:

h=W0x+αrBAx,B∈Rd×r,  A∈Rr×k,  r≪min⁡(d,k).h = W_0 x + \frac{\alpha}{r} B A x, \qquad B \in \mathbb{R}^{d \times r},\; A \in \mathbb{R}^{r \times k},\; r \ll \min(d, k).

Tiny example. A 4,096 × 4,096 projection has 16.8 million weights. A rank-8 update trains 8×(4,096+4,096)=65,5368 \times (4{,}096 + 4{,}096) = 65{,}536 numbers: 0.39% of the matrix. Adam's states are needed only for those.

LoRA: a frozen matrix plus a low-rank update
W0frozen · 4,096 × 4,096+B×A · 8 × 4,096trainable · rank 8
This matrix, full fine-tuning
16,777,216
This matrix, LoRA
65,536 (0.39%)
Llama 3 8B, query + value in all 32 layers
3.4M (0.043% of 8B)

Trainable numbers are r × (d + k) per matrix. B starts at zero, so training begins from the unchanged model. Llama 3 8B’s value projection is 4,096 × 1,024 because of grouped-query attention. Counts are our arithmetic from the published architecture; Adam’s states are needed only for the trainable part.

Compared with fine-tuning GPT-3 175B with Adam, LoRA reduced the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times, while matching or beating fine-tuning quality on the models tested Established. With rank 4 applied only to the query and value projections, the checkpoint shrank from 350 GB to about 35 MB; storing 100 adapted models takes about 354 GB instead of 35 TB Established.

No extra latency, or many tasks at once

For serving, W=W0+BAW = W_0 + BA can be computed and stored, so inference runs as usual with no added latency; switching tasks means subtracting one BABA and adding another Established. Alternatively, keep the adapters separate and let one copy of the base model serve requests for many adapters.

QLoRA: a 4-bit base

LoRA still keeps the frozen base in 16 bits. QLoRA backpropagates through a frozen 4-bit base model into LoRA adapters, which reduced memory enough to fine-tune a 65B-parameter model on a single 48 GB GPU while preserving 16-bit fine-tuning task performance Established. It introduced the 4-bit NormalFloat (NF4) data type for normally distributed weights, double quantization of the quantization constants, and paged optimizers to absorb memory spikes Established. See quantization for why the constants matter.

Why should I care?

As a researcher

That a rank-4 or rank-8 update suffices for many adaptations says the useful change from fine-tuning is low-dimensional, an observation the LoRA paper investigates.

As an engineer

LoRA and QLoRA are how most teams fine-tune open models on modest hardware, and how one base model can serve many customers with small per-customer adapters.

Modern systems that depend on it

  • QLoRA
  • adapter serving
  • instruction tuning on one GPU
  • domain adaptation

Historical context

Before

Adapting a pretrained model meant fine-tuning all its weights and storing a full copy per task, or adding adapter layers that slowed inference.

After

Frozen base model plus small trainable low-rank updates; with QLoRA, the frozen base is stored in 4 bits.

Used today

Parameter-efficient fine-tuning, mostly LoRA and its variants, is a standard option in fine-tuning libraries and services for open models.

What to remember

  • W = W₀ + (α/r)·BA; W₀ frozen, B is d × r, A is r × k.
  • B starts at zero, so training starts from the original model.
  • Trainable parameters per matrix: r(d + k) instead of dk (0.4% for a 4,096 × 4,096 matrix at r = 8).
  • Merge BA into W₀ for serving: no extra latency. Keep it separate to swap tasks.
  • QLoRA: 4-bit NF4 base + LoRA; a 65B model fine-tuned on one 48 GB GPU.

Key papers

Essential

QLoRA: Efficient Finetuning of Quantized LLMs

Tim Dettmers, Artidoro Pagnoni et al. · 2023

Combined 4-bit quantization with LoRA so a 65B model can be fine-tuned on one 48 GB GPU, putting large-model fine-tuning within reach of small labs.

How to read it: Section 3 explains NF4, double quantization and paged optimizers in two pages; the block-size arithmetic (0.5 → 0.127 bits per parameter) is worth checking by hand.

~45 min readarXiv:2305.14314✓ verified 2026-10-05
Essential

LoRA: Low-Rank Adaptation of Large Language Models

Edward J. Hu, Yelong Shen et al. · 2021

The standard way to adapt a large model cheaply: freeze its weights and learn a small low-rank update for some matrices.

How to read it: Section 4 (the method, one page) and Section 7 (why a low rank is enough) are the essentials.

~40 min readarXiv:2106.09685✓ verified 2026-10-05
Optional

Parameter-Efficient Transfer Learning for NLP

Neil Houlsby, Andrei Giurgiu et al. · 2019

Introduced adapters, small trainable layers inserted into a frozen network: the start of parameter-efficient fine-tuning.

How to read it: Figure 2 shows where the adapters go; that picture is most of the paper.

~25 min readarXiv:1902.00751✓ verified 2026-10-05

Watch