Skip to content
Road to Intelligence

Concept · Chapter 11: Inside Modern LLMs

Mixture of Experts (MoE)

Must knowKnow well15 minDifficulty

A mixture-of-experts layer replaces one feed-forward network with many smaller 'experts' and a router that sends each token to only a few of them, so the model can store far more parameters than it uses for any one token.

The problem

Quality improves with parameters, but in a dense model every parameter costs computation for every token, so capacity and cost grow together.

The solution

Keep many feed-forward experts per layer and a small learned router; for each token, run only the top-k experts by router score and add their outputs, weighted by the router's (renormalized) probabilities. Add a mechanism that keeps the experts evenly used.

The consequence

Per-token compute follows the active parameters while memory follows the total. MoE models need all experts in memory, careful load balancing and fast communication between GPUs, but many frontier and open models now use them.

Many experts, a few at a time

The feed-forward layer holds most of a Transformer's parameters. A mixture-of-experts (MoE) layer keeps several copies, each with its own weights, called experts, and adds a router: a small linear layer that scores each expert for the current token.

Mixtral's router computes G(x)=Softmax(TopK(x⋅Wg))G(x) = \text{Softmax}(\text{TopK}(x \cdot W_g)): it keeps the top-K expert scores, sets the others to minus infinity, and takes a softmax, so only K experts receive non-zero weight Established. The layer's output is

y=∑i∈top-KG(x)i Ei(x).y = \sum_{i \in \text{top-}K} G(x)_i \, E_i(x).

Tiny example. Eight experts, top-2. A token's router scores are [1.2,−0.3,2.0,0.1,0.4,−1.0,1.9,0.0][1.2, -0.3, 2.0, 0.1, 0.4, -1.0, 1.9, 0.0]. The top two are expert 3 (2.0) and expert 7 (1.9). Softmax over just those: e2.0/(e2.0+e1.9)≈0.52e^{2.0}/(e^{2.0}+e^{1.9}) \approx 0.52 and 0.480.48. The output is 0.52 E3(x)+0.48 E7(x)0.52\,E_3(x) + 0.48\,E_7(x); the other six experts are never run for this token. The next token, or the same token in the next layer, can choose different ones.

Total versus active parameters

Shazeer and colleagues reported over 1,000× improvements in model capacity with only minor losses in computational efficiency, using MoE layers of up to 137 billion parameters between LSTM layers Established. Mixtral 8x7B has 8 experts per layer and routes each token to 2; each token has access to 47B parameters but uses only 13B during inference, and Mixtral matched or outperformed Llama 2 70B and GPT-3.5 on the benchmarks reported Established.

The catch is memory. Mixtral's authors note that the memory cost of serving it is proportional to its 47B total parameters, still smaller than Llama 2 70B, and that routing adds overhead and extra memory loads Established. Every expert must sit in GPU memory even though each token uses two.

DeepSeek-V3 has 671B total parameters with 37B active per token; each MoE layer has 1 shared expert and 256 routed experts, of which 8 are activated per token Established.

Keeping the experts busy

Left alone, a router tends to favour a few experts, which then get most tokens while others idle and learn little. The Switch Transformer routes each token to a single expert and adds an auxiliary load-balancing loss, minimized when tokens are spread uniformly across experts, with a coefficient of 0.01 Established. Each expert also has a capacity, a maximum number of tokens per batch. In Switch, tokens beyond an expert's capacity are dropped from that layer: their computation is skipped and the representation passes on through the residual connection Established. With these changes Switch models pretrained up to 7× faster than T5-Base with the same compute Established. DeepSeek-V3 instead balances load without an auxiliary loss Established.

One mixture-of-experts layer · 8 experts, top-2 routing
Token:
  1. Expert 11.2idle
  2. Expert 2-0.3idle
  3. Expert 30.52runs
  4. Expert 40.1idle
  5. Expert 50.4idle
  6. Expert 6-1.0idle
  7. Expert 70.48runs
  8. Expert 80.0idle

Output = 0.52 × expert 3 + 0.48 × expert 7. Six experts do no work for this token, but all eight stay in memory. Try another token: it can choose different experts.

Mixtral 8x7B, per token

stored (memory)47B
used (compute)13B

Router scores are made up for illustration; the top-2 selection and softmax are computed. Mixtral’s parameter counts are from its paper.

The hardware reality

Experts are usually spread across GPUs (expert parallelism), so each MoE layer sends tokens to the GPUs that hold their experts and gathers the results back. That all-to-all communication, the uneven load and the memory for idle experts are the costs of the saved arithmetic. MoE pays off when compute per token is the bottleneck and there is memory and network bandwidth to spare.

Why should I care?

As a researcher

MoE separates two quantities scaling laws usually tie together, parameters and compute per token, and raises its own questions: routing, specialization, balancing and stability.

As an engineer

An MoE model's memory, speed and cost don't follow its headline parameter count. You budget memory for the total, compute for the active parameters, and network bandwidth for sending tokens to experts on other GPUs.

Modern systems that depend on it

  • Mixtral
  • DeepSeek-V3
  • expert parallelism
  • sparse scaling laws

Historical context

Before

Dense Transformers: every token passes through every parameter of every feed-forward layer.

After

Sparse layers where each token uses a few experts out of many; models are described by total and active parameters (47B and 13B for Mixtral).

Used today

Open models such as Mixtral and DeepSeek-V3 are mixture-of-experts Transformers, and several frontier labs have described sparse models.

What to remember

  • Router: a linear layer giving one score per expert; keep the top-k, softmax over them, mix those experts' outputs.
  • Total parameters set memory; active parameters set compute per token.
  • Mixtral: 8 experts per layer, top-2, 47B total, 13B active.
  • DeepSeek-V3: 256 routed experts + 1 shared, 8 routed active, 671B total, 37B active.
  • Without balancing, a few experts get most tokens; Switch adds a balancing loss and a capacity limit.

Key papers

Essential

Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Noam Shazeer, Azalia Mirhoseini et al. · 2017

Made conditional computation work at scale: a learned gate sends each input to a few of many expert networks, so parameters can grow far faster than compute.

How to read it: Section 2 (the gating network) and Section 4 (balancing expert use) are the parts still in use today.

~40 min readarXiv:1701.06538✓ verified 2026-10-05
Important

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

William Fedus, Barret Zoph, Noam Shazeer · 2021

Simplified mixture-of-experts Transformers (route each token to one expert) and made them train stably, including in bfloat16.

How to read it: Section 2 (routing, capacity factor and the balancing loss in Equation 4) is the core; the rest is scaling experiments.

~1 h readarXiv:2101.03961✓ verified 2026-10-05
Important

Mixtral of Experts

Albert Q. Jiang, Alexandre Sablayrolles et al. · 2024

A widely used open mixture-of-experts LLM, and a clean example of total versus active parameters: 47B stored, 13B used per token.

How to read it: Section 2 is a page long and covers the whole architecture; Section 5 looks at which experts tokens actually choose.

~20 min readarXiv:2401.04088✓ verified 2026-10-05
Important

DeepSeek-V3 Technical Report

DeepSeek-AI et al. · 2024

A detailed report on a frontier-scale open mixture-of-experts model that brings this chapter's ideas together: sparse experts, a compressed attention cache, FP8 training.

How to read it: Read Section 2 (architecture) first; Section 3 on infrastructure is valuable if you work on systems.

~1 h 30 min readarXiv:2412.19437✓ verified 2026-10-05

Watch