Concept · Chapter 11: Inside Modern LLMs
The Modern Decoder Block: Pre-Norm, RMSNorm, SwiGLU
Today's open LLMs keep the 2017 Transformer's decoder block but normalize before each sublayer instead of after (pre-norm), use the cheaper RMSNorm, and replace the feed-forward layer's single nonlinearity with a gated one (SwiGLU).
The problem
The original block trained less stably without a careful warmup, spent time re-centring every vector, and used a plain ReLU feed-forward layer that later experiments could beat.
The solution
Put normalization inside each residual branch so the residual path stays clean; drop LayerNorm's mean subtraction (RMSNorm); multiply one projection by a gated second projection (SwiGLU).
The consequence
These small changes, plus rotary positions and grouped-query attention, became a shared default ('LLaMA-style'). None changes what a block does in principle; they make it train more reliably and run a little faster.
You should understand first
Same block, three edits
The Transformer block of Chapter 7 is still there: attention, then a feed-forward layer, each wrapped in a residual connection with normalization. Llama 2 describes its architecture as the standard Transformer with pre-normalization using RMSNorm, the SwiGLU activation and rotary positional embeddings Established, which is a fair summary of what most open models changed.
1. Normalize first (pre-norm)
The 2017 block normalized after adding the residual: . Most models now normalize inside the branch: . The residual path from input to output is then an untouched sum.
Xiong and colleagues showed that at initialization the post-norm Transformer has large gradients near the output, which is why it needs a learning-rate warmup, while pre-norm's gradients are well-behaved and it can train without warmup Established. Large runs still warm up (learning-rate schedules), but they are less fragile.
2. RMSNorm: skip the mean
LayerNorm subtracts each vector's mean and divides by its standard deviation. RMSNorm only divides by the root mean square:
Tiny example. . LayerNorm: mean 1, standard deviation 1.58, giving . RMSNorm: root mean square , giving . Both bring the vector to a standard scale; only LayerNorm also centres it.
Zhang and Sennrich hypothesized that LayerNorm's re-centring is dispensable; RMSNorm matched LayerNorm's quality while reducing running time by 7–64% across the models they tested Established.
3. A gated feed-forward layer (SwiGLU)
The original feed-forward layer is . A gated linear unit multiplies one projection by an activated copy of another:
Shazeer tested several such variants in the Transformer's feed-forward layer, shrinking the hidden width so the parameter count stayed constant, and found some of them, SwiGLU among them, improved quality over ReLU and GELU Established. He offered no theory for why. Llama 3 8B uses SwiGLU with a model dimension of 4,096 and a feed-forward dimension of 14,336 Established.
What to remember
- Pre-norm: x + Sublayer(Norm(x)). Post-norm (2017): Norm(x + Sublayer(x)).
- Pre-norm keeps gradients well-behaved at initialization; warmup matters less.
- RMSNorm divides by the root mean square, with no mean subtraction; 7–64% faster than LayerNorm in its paper.
- SwiGLU: (SiLU(xW) ⊙ xV) W₂, three matrices instead of two.
- Llama 2 and 3: pre-normalization with RMSNorm, SwiGLU, RoPE (plus GQA).
Key papers
On Layer Normalization in the Transformer Architecture
Ruibin Xiong, Yunchang Yang et al. · 2020 · ICML 2020
Explains why modern Transformers put layer normalization before each sublayer ('pre-LN') rather than after it.
Root Mean Square Layer Normalization
Biao Zhang, Rico Sennrich · 2019
RMSNorm, LayerNorm without the mean subtraction, is the normalization in LLaMA-style models.
How to read it: Equations 2–4 are the whole method; the rest is experiments.
GLU Variants Improve Transformer
Noam Shazeer · 2020
The origin of SwiGLU, the gated feed-forward layer used by LLaMA, Mistral and many others.
How to read it: Four pages. Equations 5–6 define the variants; note the author's own admission that the paper offers results, not an explanation.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin et al. · 2023
A widely used open model family; its larger models adopted grouped-query attention for faster inference.
How to read it: For this chapter, read Section 2.2 and Appendix A.2.1, which compares multi-head, multi-query and grouped-query attention.
Watch
Stanford Online
Stanford CS336 Lang. Modeling from Scratch | Spring 2025 | Lec. 3: Architectures, Hyperparameters
A survey of what changed inside the Transformer block between 2017 and today's open models, and which choices nearly everyone now agrees on.