Concept · Chapter 12: Embeddings, RAG & the LLM Application Stack
Tool Calling
Tool calling lets a model ask the application to run a function: the model emits a structured request (a tool name and JSON arguments), the application executes it and returns the result as a new message, and the model continues with that result in its context.
The problem
Language models are bad at things simple software does well (arithmetic, looking up current data, querying a database, taking actions), and text alone can't do anything in the world.
The solution
Describe the available tools (name, purpose, argument schema) in the context; train or prompt the model to emit a call when one is needed; let the application run it, append the result, and generate again.
The consequence
Models can use calculators, search, databases and APIs, which made assistants useful for live data and actions. It also made them a security concern: the model's choices now have effects, and tool results are untrusted input. Running calls in a loop is what Chapter 13 calls an agent.
You should understand first
- Text as Data
- Probability and Distributions
- Conditional Probability and Bayes' Theorem
- Probability of Sequences
- Language Modeling
- Vectors
- Dot Product
- Embeddings
- Attention
- Softmax
- Self-Attention
- Causal Masking
- Entropy
- Loss Functions
- Cross-Entropy Loss
- Autoregressive Next-Token Prediction
- Decoding: Greedy, Temperature, Top-k, Top-p
- One-Hot Encoding
- Tokenization
- Chat Templates
- Pretraining at Scale
- GPT-1 → GPT-2 → GPT-3
- In-Context Learning
- System Prompts and Instructions
- Structured Outputs and Constrained Decoding
- Tool Calling
The round trip
- The application sends the conversation plus tool definitions: a name, a description and a JSON schema for the arguments, for example
get_weather(city: string, date: string). - The model, instead of answering in prose, emits a call:
{"name": "get_weather", "arguments": {"city": "Kathmandu", "date": "2026-10-06"}}. This is structured output, often marked by special tokens in the chat template. - The application validates the arguments, checks the user is allowed to do this, and runs the function.
- The result is appended as a tool message:
{"temp_c": 22, "rain_chance": 0.6}. - The model generates again, now able to read the result: "Expect about 22 °C with a 60% chance of rain."
The model never executes anything. In the notes under his introduction-to-LLMs video, Karpathy describes the mechanism: the model emits special words, the program running it detects them, sends the request to the tool, and continues generation with the result; fine-tuning data teaches when and how to call tools Established.
How models learn it
Toolformer taught a model to call a calculator, a question-answering system, search engines, a translation system and a calendar in a self-supervised way, keeping only the calls whose results helped predict the following text Established. Commercial APIs followed: OpenAI announced function calling on 13 June 2023, letting developers describe functions so the model can choose to output a JSON object containing arguments to call them Established. Today's models are post-trained on many examples of tool definitions, calls and results.
Retrieval is a tool too
A RAG pipeline that searches on every question is fixed. Expose search as a tool instead and the model decides when to search and what query to send, and can search again if the results are poor. That's a step towards agents.
Effects and risks
Once a model can call send_email or delete_file, its mistakes have consequences. And the results it reads can be hostile: OpenAI's announcement itself noted a proof-of-concept exploit in which untrusted data from a tool's output instructs the model to perform unintended actions Established. The defences are ordinary engineering: least privilege per tool, validation, confirmation before irreversible actions, and treating tool output as data (guardrails).
One call, then a loop
Everything here is one round trip. Let the model call tools repeatedly, reading each result and deciding what to do next until the task is done, and you have an agent, the subject of Chapter 13. The Model Context Protocol, which Anthropic open-sourced in November 2024 Established, standardises how tools and data sources are offered to models, so one integration works across applications.
Why should I care?
As a researcher
Tool use changes what 'model capability' means: a model plus a calculator and a search engine is a different system to evaluate, and learning when to call a tool is itself a training problem.
As an engineer
Tool calling is how a model connects to your systems; the schema design, validation and permission checks around each tool are where reliability and safety are won or lost.
Modern systems that depend on it
- agents (Chapter 13)
- search-augmented chat
- code execution
- the Model Context Protocol
Historical context
Before
Developers parsed free text for commands, or wrote rigid pipelines where the model filled a template.
After
The model chooses among declared tools and fills typed arguments; the application keeps control of execution.
Used today
Major model APIs support declaring tools with JSON schemas; chat assistants use it for web search, code execution and file access.
What to remember
- The model never runs anything: it asks; the application executes.
- A call is structured output: tool name + arguments matching a schema.
- The result goes back into the context as a message, then the model continues.
- Validate arguments and check permissions in code before executing.
- Tool results are untrusted text: they can contain injected instructions.
Key papers
Toolformer: Language Models Can Teach Themselves to Use Tools
Timo Schick, Jane Dwivedi-Yu et al. · 2023
Showed a model can learn when to call a calculator, search engine or calendar and how to use the result, the core idea behind tool calling.
Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Kai Greshake, Sahar Abdelnabi et al. · 2023
Showed that text a model retrieves (a web page, an email, a document) can carry instructions an attacker planted, so retrieval and tools are a security boundary.
Watch
Andrej Karpathy
[1hr Talk] Intro to Large Language Models
A clear one-hour overview of what LLMs are, how they are trained, and where they're going — good orientation for Part III.
Andrej Karpathy
Deep Dive into LLMs like ChatGPT
A long, general-audience walk through the whole pipeline behind a chat model, from internet text to tokens to pretraining to post-training.