Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Structured Outputs and Constrained Decoding
Structured outputs make a model produce data a program can parse, such as JSON matching a schema, and constrained decoding guarantees it by masking, at every step, the tokens that would break the format and renormalising over the rest.
The problem
Software needs exact formats, but a sampled model output can include a stray sentence, a missing quote or a wrong field name, and one malformed reply breaks the program that reads it.
The solution
Describe the format (instructions, examples, a JSON schema); for a guarantee, constrain decoding: track where in the format the output is, allow only tokens that keep it valid, and sample from the renormalised remainder.
The consequence
Parsing failures can be ruled out entirely, which makes model outputs safe to feed into code and tool calls. The guarantee covers syntax only: a valid object can still contain a wrong or invented value.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
Text versus data
A person can read "Sure! Here's the JSON: {name: ..." and shrug. A program calling JSON.parse crashes. Applications need replies they can parse: extracted fields, classifications, tool arguments. Instructions and examples get you most of the way; models are trained to follow formats, but sampling means a small fraction of replies still break.
Constrained decoding
Decoding already edits the model's next-token distribution: top-k and top-p zero out some tokens and renormalise the rest. Constrained decoding does the same with a different rule: zero out every token that would make the output impossible to complete validly.
Tiny example. The output so far is {"city": "Kathmandu" and the model's next-token probabilities are , 0.50, } 0.30, and 0.15, ] 0.05. Only , and } can continue a valid object here, so the constrained distribution is , 0.50/0.80 = 0.625, } 0.30/0.80 = 0.375. Every sample now stays valid; the model's preferences between valid options are kept.
The hard part is knowing which tokens are allowed, since tokens don't line up with JSON's characters. Willard and Louf compile a regular expression or grammar into a finite-state machine and index the vocabulary against its states in advance, so the set of allowed tokens costs O(1) on average to obtain at each step instead of a check over the whole vocabulary Established. Libraries and model APIs now offer schema-constrained generation; OpenAI's API changelog lists Structured Outputs, model outputs that reliably adhere to developer-supplied JSON schemas, as launched on 6 August 2024 Established.
Valid is not correct
The constraint guarantees the shape, not the content. If the schema requires a date and the request didn't contain one, constrained decoding will still produce a date, because "no date" isn't an allowed continuation. The model is forced to invent. The fix is in the schema: make fields optional or nullable when "unknown" is an honest answer, and validate values (ranges, enums, cross-checks) in code afterwards. This chapter's lab shows both: invalid outputs disappear, and an invented date appears.
What to remember
- Asking nicely for JSON usually works and sometimes doesn't.
- Constrained decoding: mask invalid tokens, renormalise, sample. Same move as top-k and top-p.
- A finite-state machine or grammar tracks which tokens are allowed next.
- Valid ≠ correct: the schema can force a value the model doesn't know.
- Let the schema say 'unknown' (null, optional fields) when that's an honest answer.
Key papers
Efficient Guided Generation for Large Language Models
Brandon T. Willard, Rémi Louf · 2023
Made constrained decoding cheap: precompute which tokens are allowed in each state of a regular expression or grammar, so structured output can be guaranteed.
How to read it: The section on iterative FSM processing and indexing, with Figure 1's floating-point-number example, is the core.