Concept · Chapter 10: From Base Model to Assistant
Chat Templates
A chat template serializes roles and messages into the token sequence a particular model expects.
The problem
An API conversation is a list of structured messages, but a causal language model consumes one token sequence.
The solution
Apply the model's template to mark roles, message boundaries and the position where the assistant should continue.
The consequence
Correct formatting connects the application to the model's training distribution; wrong markers can damage otherwise good responses.
You should understand first
- Text as Data
- Vectors
- One-Hot Encoding
- Tokenization
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Chat Templates
From messages to a sequence
An application might hold three records: system guidance, a user question and an assistant answer. The model sees tokens. A template chooses how to join the records and which special tokens mark each boundary.
For a schematic, not real model syntax, imagine:
START · USER · question · END-TURN · ASSISTANT · answer · END-TURN
Different models use different delimiters and may handle system messages differently. Use the checkpoint's actual template rather than borrowing markers from a familiar model. The Transformers chat-template guide shows why this matters.
Training and generation meet here
During SFT, the serialized sequence includes the demonstration answer. The loss mask selects which tokens are targets. During generation, a template may append the assistant prefix so the model knows where to continue. Continuing an unfinished assistant message is a different operation from starting a new one.
Avoid adding special beginning/end tokens twice when formatting and tokenizing separately. Test a rendered example before scaling up a dataset: inspect role markers, end tokens and the loss mask together.
Structure is not enforcement
A system-role marker tells the model which learned convention to apply. It does not mathematically prevent a later instruction from changing behavior. Prompt injection and application security need a separate treatment in the systems and safety chapters.
Check yourself: if training ends every assistant response with a turn terminator but inference starts with an unexpected delimiter, what changed? The weights did not change; the input format moved away from the one those weights learned.
What to remember
- Templates are model-specific; role names alone do not specify their tokens.
- Training includes target responses; inference may need an assistant generation prefix.
- A role marker is not a complete security boundary.
Key papers
The Llama 3 Herd of Models
Aaron Grattafiori, Abhimanyu Dubey et al. · 2024
The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.
How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.
Watch
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.