Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

The Modern Decoder Block: Pre-Norm, RMSNorm, SwiGLU

Should knowUnderstand11 minDifficulty

Today's open LLMs keep the 2017 Transformer's decoder block but normalize before each sublayer instead of after (pre-norm), use the cheaper RMSNorm, and replace the feed-forward layer's single nonlinearity with a gated one (SwiGLU).

The problem

The original block trained less stably without a careful warmup, spent time re-centring every vector, and used a plain ReLU feed-forward layer that later experiments could beat.

The solution

Put normalization inside each residual branch so the residual path stays clean; drop LayerNorm's mean subtraction (RMSNorm); multiply one projection by a gated second projection (SwiGLU).

The consequence

These small changes, plus rotary positions and grouped-query attention, became a shared default ('LLaMA-style'). None changes what a block does in principle; they make it train more reliably and run a little faster.

Same block, three edits

The Transformer block of Chapter 7 is still there: attention, then a feed-forward layer, each wrapped in a residual connection with normalization. Llama 2 describes its architecture as the standard Transformer with pre-normalization using RMSNorm, the SwiGLU activation and rotary positional embeddings Established, which is a fair summary of what most open models changed.

1. Normalize first (pre-norm)

The 2017 block normalized after adding the residual: Norm(x+Sublayer(x))\text{Norm}(x + \text{Sublayer}(x)). Most models now normalize inside the branch: x+Sublayer(Norm(x))x + \text{Sublayer}(\text{Norm}(x)). The residual path from input to output is then an untouched sum.

Xiong and colleagues showed that at initialization the post-norm Transformer has large gradients near the output, which is why it needs a learning-rate warmup, while pre-norm's gradients are well-behaved and it can train without warmup Established. Large runs still warm up (learning-rate schedules), but they are less fragile.

2. RMSNorm: skip the mean

LayerNorm subtracts each vector's mean and divides by its standard deviation. RMSNorm only divides by the root mean square:

RMSNorm(a)i=ai1n∑jaj2 gi.\text{RMSNorm}(a)_i = \frac{a_i}{\sqrt{\tfrac{1}{n}\sum_j a_j^2}} \, g_i .

Tiny example. a=[2,−1,3,0]a = [2, -1, 3, 0]. LayerNorm: mean 1, standard deviation 1.58, giving [0.63,−1.26,1.26,−0.63][0.63, -1.26, 1.26, -0.63]. RMSNorm: root mean square 14/4=1.87\sqrt{14/4} = 1.87, giving [1.07,−0.53,1.60,0][1.07, -0.53, 1.60, 0]. Both bring the vector to a standard scale; only LayerNorm also centres it.

Zhang and Sennrich hypothesized that LayerNorm's re-centring is dispensable; RMSNorm matched LayerNorm's quality while reducing running time by 7–64% across the models they tested Established.

3. A gated feed-forward layer (SwiGLU)

The original feed-forward layer is max⁡(0,xW1)W2\max(0, xW_1)W_2. A gated linear unit multiplies one projection by an activated copy of another:

FFNSwiGLU(x)=(SiLU(xW)⊙xV)W2,SiLU(z)=z σ(z).\text{FFN}_{\text{SwiGLU}}(x) = \big(\text{SiLU}(xW) \odot xV\big) W_2, \qquad \text{SiLU}(z) = z\,\sigma(z).

Shazeer tested several such variants in the Transformer's feed-forward layer, shrinking the hidden width so the parameter count stayed constant, and found some of them, SwiGLU among them, improved quality over ReLU and GELU Established. He offered no theory for why. Llama 3 8B uses SwiGLU with a model dimension of 4,096 and a feed-forward dimension of 14,336 Established.

What to remember

  • Pre-norm: x + Sublayer(Norm(x)). Post-norm (2017): Norm(x + Sublayer(x)).
  • Pre-norm keeps gradients well-behaved at initialization; warmup matters less.
  • RMSNorm divides by the root mean square, with no mean subtraction; 7–64% faster than LayerNorm in its paper.
  • SwiGLU: (SiLU(xW) ⊙ xV) W₂, three matrices instead of two.
  • Llama 2 and 3: pre-normalization with RMSNorm, SwiGLU, RoPE (plus GQA).

Key papers

Optional

On Layer Normalization in the Transformer Architecture

Ruibin Xiong, Yunchang Yang et al. · 2020 · ICML 2020

Explains why modern Transformers put layer normalization before each sublayer ('pre-LN') rather than after it.

~1 h readarXiv:2002.04745✓ verified 2026-09-26
Optional

Root Mean Square Layer Normalization

Biao Zhang, Rico Sennrich · 2019

RMSNorm, LayerNorm without the mean subtraction, is the normalization in LLaMA-style models.

How to read it: Equations 2–4 are the whole method; the rest is experiments.

~20 min readarXiv:1910.07467✓ verified 2026-10-05
Optional

GLU Variants Improve Transformer

Noam Shazeer · 2020

The origin of SwiGLU, the gated feed-forward layer used by LLaMA, Mistral and many others.

How to read it: Four pages. Equations 5–6 define the variants; note the author's own admission that the paper offers results, not an explanation.

~10 min readarXiv:2002.05202✓ verified 2026-10-05
Important

Llama 2: Open Foundation and Fine-Tuned Chat Models

Hugo Touvron, Louis Martin et al. · 2023

A widely used open model family; its larger models adopted grouped-query attention for faster inference.

How to read it: For this chapter, read Section 2.2 and Appendix A.2.1, which compares multi-head, multi-query and grouped-query attention.

~1 h 30 min readarXiv:2307.09288✓ verified 2026-10-05

Watch