Skip to content
Road to Intelligence

Concept · Chapter 10: From Base Model to Assistant

Chat Templates

Must knowUnderstand7 minDifficulty

A chat template serializes roles and messages into the token sequence a particular model expects.

The problem

An API conversation is a list of structured messages, but a causal language model consumes one token sequence.

The solution

Apply the model's template to mark roles, message boundaries and the position where the assistant should continue.

The consequence

Correct formatting connects the application to the model's training distribution; wrong markers can damage otherwise good responses.

From messages to a sequence

An application might hold three records: system guidance, a user question and an assistant answer. The model sees tokens. A template chooses how to join the records and which special tokens mark each boundary.

For a schematic, not real model syntax, imagine:

START · USER · question · END-TURN · ASSISTANT · answer · END-TURN

Different models use different delimiters and may handle system messages differently. Use the checkpoint's actual template rather than borrowing markers from a familiar model. The Transformers chat-template guide shows why this matters.

Training and generation meet here

During SFT, the serialized sequence includes the demonstration answer. The loss mask selects which tokens are targets. During generation, a template may append the assistant prefix so the model knows where to continue. Continuing an unfinished assistant message is a different operation from starting a new one.

Avoid adding special beginning/end tokens twice when formatting and tokenizing separately. Test a rendered example before scaling up a dataset: inspect role markers, end tokens and the loss mask together.

Structure is not enforcement

A system-role marker tells the model which learned convention to apply. It does not mathematically prevent a later instruction from changing behavior. Prompt injection and application security need a separate treatment in the systems and safety chapters.

Check yourself: if training ends every assistant response with a turn terminator but inference starts with an unexpected delimiter, what changed? The weights did not change; the input format moved away from the one those weights learned.

What to remember

  • Templates are model-specific; role names alone do not specify their tokens.
  • Training includes target responses; inference may need an assistant generation prefix.
  • A role marker is not a complete security boundary.

Key papers

Essential

The Llama 3 Herd of Models

Aaron Grattafiori, Abhimanyu Dubey et al. · 2024

The most complete public account of building a frontier-scale model end to end: data pipeline, scaling-law experiments, 16,384-GPU training, failures and all.

How to read it: It is 90+ pages. For this chapter read Section 3 (pre-training) only: data, scaling laws, infrastructure and the training recipe.

~2 h readarXiv:2407.21783✓ verified 2026-10-04

Watch

3 h 31 min

Andrej Karpathy

Deep Dive into LLMs like ChatGPT

A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.

Should know