Tool-calling agent¶
An agent whose tools reach external systems: the model decides which to call, and when.
Tools are ordinary Python functions you register on an agent. The model sees each tool's name, arguments and docstring, decides when a call would help, and the result goes back into the conversation so it can decide what to do next. The loop repeats until it has an answer. Most of what makes tools reliable is in how you write them: keep the interface simple, report errors in plain English the model can act on, and keep large payloads inside the tool and out of the context. This example is a lookup over Python release notes, with the backend behind the dependencies so you can swap in an API client or a database.
Use it when
- The agent has to read from or act on the outside world: an API, a database, files.
- What the model should do next depends on what a tool returned.
- You want the model to choose which tool to call, and when, and not follow a fixed sequence.
Look elsewhere when
- The tools already exist as an MCP server shared with other clients:
mcp_tools. - One question needs dozens of calls and exact arithmetic:
code_mode. - A tool can do something you can't undo: add
human_in_the_loop.
What it shows
- A working tool,
python_release_notes, over an in-memory table of Python releases. The data lives behindToolAgentDeps, which is where an API client or database connection goes - Simple tool interfaces, English error messages, large payloads held at the tool layer
- The three-outcome error convention in one tool: success,
ModelRetry(a malformed version like"latest"or"3.13.1"— the model can fix it) andToolFailed(an unknown version — nothing to find, so no retry budget is spent) agent/tools/example.pyfor the same convention in a standalone tool
See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.
To adapt it, replace RELEASES and python_release_notes with your own lookup, and keep the
three outcomes.
Source¶
All of it is in examples/tool_calling/.
"""Tool-calling agent pattern.
Use this pattern when:
- The agent needs to interact with external systems (APIs, databases, files)
- You want the LLM to decide which tools to call and when
- Tool results inform subsequent decisions (agentic loop)
This example answers questions about Python releases by calling a lookup tool. The data
lives in a small in-memory table so it runs anywhere; in a real agent, `ReleaseNotes` is
where an API client or database connection goes.
Key design principles (from production experience):
- Keep tool interfaces simple: fewer optional params = more reliable tool selection
- Translate errors into English: give the LLM enough context to self-correct
- Hold large payloads at the tool layer: don't dump raw API responses into context
- Inject the backend through deps, so tests (and you) can swap it
The tool shows all three error outcomes: success, `ModelRetry` (the model can fix its input),
and `ToolFailed` (expected and terminal — there is nothing to find). See
agent/tools/example.py for the same convention in a standalone tool.
"""
from __future__ import annotations
import re
from dataclasses import dataclass, field
from pydantic import BaseModel
from pydantic_ai import Agent, ModelRetry, RunContext, ToolFailed
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits
from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.runs import Flow, RunResult
logger = get_logger(__name__)
LABEL = agent_label(__name__) # names this agent's run spans in Logfire traces
# Guardrail against runaway agentic loops. A run that exceeds any limit
# raises UsageLimitExceeded instead of silently burning tokens. Tune per task:
# request_limit caps model round-trips (each tool-call iteration is one
# request), total_tokens_limit caps overall tokens for the run. Set
# AGENT_COST_LIMIT (USD) to add a spend cap — optional, off by default, and only
# useful for models with known pricing (see Settings.cost_limit).
USAGE_LIMITS = UsageLimits(
request_limit=10, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)
# --- The backend the tool talks to ---
@dataclass(frozen=True)
class Release:
released: str
highlights: tuple[str, ...]
# Replace with a real data source. A dict stands in for an API or database here.
RELEASES: dict[str, Release] = {
"3.11": Release(
"October 24, 2022",
(
"CPython runs about 25% faster on average (the Faster CPython project)",
"exception groups and except* (PEP 654)",
"tomllib adds TOML parsing to the standard library (PEP 680)",
),
),
"3.12": Release(
"October 2, 2023",
(
"cleaner generics syntax: type parameters and the type statement (PEP 695)",
"f-strings can nest quotes and span lines (PEP 701)",
"per-interpreter GIL for subinterpreters (PEP 684)",
),
),
"3.13": Release(
"October 7, 2024",
(
"an experimental free-threaded build with the GIL disabled (PEP 703)",
"an experimental JIT compiler (PEP 744)",
"a new interactive interpreter with multi-line editing and colour",
),
),
}
# --- Dependencies ---
@dataclass
class ToolAgentDeps:
"""Runtime dependencies for the tool-calling agent."""
# The release table the tool reads. Swap in an API client or database in a real agent.
releases: dict[str, Release] = field(default_factory=lambda: RELEASES)
# --- Output type ---
class ToolAgentOutput(BaseModel):
# `result` is the conventional output field in these examples; the generated
# eval starter reads it when present (see evals/helpers.py).
result: str
# Pydantic deep-copies mutable defaults, so a plain [] is safe here.
# Do NOT use dataclasses.field() inside a BaseModel — it is not a
# Pydantic construct (use pydantic.Field(default_factory=...) if needed).
versions: list[str] = []
# --- Agent ---
tool_agent: Agent[ToolAgentDeps, ToolAgentOutput] = Agent(
settings.model,
name=LABEL, # labels this agent's run span in Logfire traces
output_type=ToolAgentOutput,
deps_type=ToolAgentDeps,
# Fail fast when the provider filters a response, instead of retrying a
# refused request or returning partial text.
capabilities=[RaiseContentFilterError()],
instructions="""You answer questions about Python releases.
Use the python_release_notes tool for every fact about a release; do not answer from
memory. If the tool reports that nothing is available for a version, tell the user that
plainly and do not guess. In `versions`, list the versions your answer draws on.
""",
)
VERSION_FORMAT = re.compile(r"\d+\.\d+")
# --- Tools ---
@tool_agent.tool
async def python_release_notes(ctx: RunContext[ToolAgentDeps], version: str) -> str:
"""Look up the release date and headline features of a Python release.
Args:
version: A major.minor version such as "3.13". Not "3.13.1" and not "latest".
Returns:
The release date and highlights as a short paragraph.
Raises:
ModelRetry: When the version isn't in major.minor form, so the model can correct it.
ToolFailed: When there are no notes for that version (a terminal, expected failure).
"""
version = version.strip()
logger.info("Tool called", extra={"tool": "python_release_notes", "version": version})
if not VERSION_FORMAT.fullmatch(version):
# The model can fix this by changing its input, so ask it to retry.
raise ModelRetry(
f"'{version}' is not a major.minor version like '3.13'. "
"Call the tool again with just the major and minor numbers."
)
release = ctx.deps.releases.get(version)
if release is None:
# Nothing exists to find, and retrying won't change that. ToolFailed shows the model
# the failure without spending retry budget, and tells it what to do instead.
known = ", ".join(sorted(ctx.deps.releases))
raise ToolFailed(
f"There are no release notes for Python {version}. Known versions: {known}. "
"Tell the user this version is unavailable; do not guess its features."
)
highlights = "; ".join(release.highlights)
return f"Python {version} was released on {release.released}. Highlights: {highlights}."
async def run_tool_agent(
user_input: str, deps: ToolAgentDeps | None = None
) -> RunResult[ToolAgentOutput]:
"""Run the tool-calling agent.
Args:
user_input: The user's message or task description.
deps: Runtime dependencies. Created with defaults if not provided.
Returns:
A RunResult: `.output` is the validated ToolAgentOutput; the tool calls the agent made
are in `.all_messages()`.
"""
if deps is None:
deps = ToolAgentDeps()
logger.info("Running tool-calling agent", extra={"user_input": user_input})
flow = Flow(USAGE_LIMITS)
result = await flow.run(tool_agent, user_input, deps=deps)
return flow.finish(result.output)
if __name__ == "__main__":
import asyncio
configure_logging()
result = asyncio.run(run_tool_agent("What changed in Python 3.13 compared with 3.12?"))
print(result.output)
title = "Tool-calling agent"
pattern = "tool_calling"
summary = "An agent that calls tools against external systems, with the three-outcome error convention."
smoke_input = "What changed in Python 3.13 compared with 3.12?"
# A real model must use the tool to answer this (checked by the release check).
expected_tools = ["python_release_notes"]
[entrypoint]
deps = "ToolAgentDeps"
run = "run_tool_agent"
"""The release-notes tool's three outcomes, directly and through the agent loop."""
import pytest
from pydantic_ai import ModelRetry, RunContext, ToolFailed
from pydantic_ai.messages import (
ModelResponse,
RetryPromptPart,
ToolCallPart,
ToolReturnPart,
)
from pydantic_ai.models.function import AgentInfo, FunctionModel
from pydantic_ai.models.test import TestModel
from pydantic_ai.usage import RunUsage
from examples.tool_calling.agent import (
RELEASES,
Release,
ToolAgentDeps,
python_release_notes,
run_tool_agent,
tool_agent,
)
def ctx(deps: ToolAgentDeps | None = None) -> RunContext[ToolAgentDeps]:
return RunContext(deps=deps or ToolAgentDeps(), model=TestModel(), usage=RunUsage())
# --- The tool, called directly ---
async def test_a_known_version_returns_its_date_and_highlights():
notes = await python_release_notes(ctx(), "3.13")
assert "October 7, 2024" in notes
assert "free-threaded" in notes and "JIT" in notes
async def test_surrounding_whitespace_is_tolerated():
assert await python_release_notes(ctx(), " 3.12 ") == await python_release_notes(ctx(), "3.12")
@pytest.mark.parametrize("version", ["latest", "3.13.1", "", "three.thirteen", "3"])
async def test_a_malformed_version_asks_the_model_to_retry(version):
with pytest.raises(ModelRetry, match="major.minor"):
await python_release_notes(ctx(), version)
async def test_an_unknown_version_is_a_terminal_failure_that_lists_what_exists():
with pytest.raises(ToolFailed) as failure:
await python_release_notes(ctx(), "3.99")
message = str(failure.value)
assert "3.99" in message and "3.11, 3.12, 3.13" in message and "do not guess" in message
async def test_the_backend_comes_from_deps():
deps = ToolAgentDeps(releases={"9.9": Release("tomorrow", ("flying cars",))})
assert "flying cars" in await python_release_notes(ctx(deps), "9.9")
with pytest.raises(ToolFailed):
await python_release_notes(ctx(deps), "3.13") # not in this backend
def test_the_default_backend_is_the_shared_table():
assert ToolAgentDeps().releases is RELEASES
# --- Through the agent loop, with a scripted model ---
def scripted(*steps: dict):
"""A model that follows `steps` in order. A step is a tool call or the final output."""
remaining = list(steps)
def model_fn(messages, info: AgentInfo) -> ModelResponse:
step = remaining.pop(0)
if "tool" in step:
return ModelResponse(parts=[ToolCallPart(step["tool"], step["args"])])
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, step["output"])])
return FunctionModel(model_fn)
def tool_results(result) -> list[ToolReturnPart]:
return [p for m in result.all_messages() for p in m.parts if isinstance(p, ToolReturnPart)]
async def test_the_model_calls_the_tool_and_answers_from_its_result():
model = scripted(
{"tool": "python_release_notes", "args": {"version": "3.13"}},
{"output": {"result": "3.13 added a JIT.", "versions": ["3.13"]}},
)
with tool_agent.override(model=model):
result = await run_tool_agent("What's new in 3.13?")
assert result.output.versions == ["3.13"]
tool = [r for r in tool_results(result) if r.tool_name == "python_release_notes"][0]
assert "October 7, 2024" in tool.content and tool.outcome == "success"
async def test_a_malformed_call_is_corrected_on_retry():
model = scripted(
{"tool": "python_release_notes", "args": {"version": "latest"}},
{"tool": "python_release_notes", "args": {"version": "3.13"}},
{"output": {"result": "ok", "versions": ["3.13"]}},
)
with tool_agent.override(model=model):
result = await run_tool_agent("What's the latest Python?")
retries = [p for m in result.all_messages() for p in m.parts if isinstance(p, RetryPromptPart)]
assert len(retries) == 1 and "major.minor" in str(retries[0].content)
assert result.output.result == "ok"
async def test_an_unavailable_version_reaches_the_model_as_a_failure_not_a_retry():
model = scripted(
{"tool": "python_release_notes", "args": {"version": "3.99"}},
{"output": {"result": "No notes exist for 3.99.", "versions": []}},
)
with tool_agent.override(model=model):
result = await run_tool_agent("What's new in 3.99?")
tool = [r for r in tool_results(result) if r.tool_name == "python_release_notes"][0]
assert tool.outcome == "failed" and "do not guess" in tool.content
assert not [p for m in result.all_messages() for p in m.parts if isinstance(p, RetryPromptPart)]
assert result.usage.requests == 2
"""Live check: the model uses the tool, answers from it, and handles a failure. `pytest -m eval`."""
import re
import pytest
from pydantic_ai.messages import ToolReturnPart
from evals.trace import traced_run
from examples.live_support import assert_every_agent_ran, run_as_script
from examples.tool_calling import agent as module
pytestmark = pytest.mark.eval
# Everything the table says about the versions it does know. An answer about a version it has no
# notes for must contain none of it (that would be borrowing another release's facts).
KNOWN_FACTS = {r.released.lower() for r in module.RELEASES.values()} | {
pep.lower()
for r in module.RELEASES.values()
for highlight in r.highlights
for pep in re.findall(r"PEP \d+", highlight)
}
def tool_returns(traced) -> list[ToolReturnPart]:
return [
p
for m in traced.result.all_messages()
for p in m.parts
if isinstance(p, ToolReturnPart) and p.tool_name == "python_release_notes"
]
@pytest.fixture(scope="module")
async def known():
return await traced_run(
module.run_tool_agent, "What changed in Python 3.13 compared with 3.12?"
)
@pytest.fixture(scope="module")
async def unknown():
return await traced_run(module.run_tool_agent, "What's new in Python 3.99?")
async def test_a_known_version_is_answered_from_the_tool(known):
assert known.tools_called and set(known.tools_called) == {"python_release_notes"}
assert {r.outcome for r in tool_returns(known)} == {"success"}
answer = known.result.output.result.lower()
assert "october 7, 2024" in answer or "free-threaded" in answer or "jit" in answer
assert "3.13" in known.result.output.versions
assert known.result.usage.requests >= 2 # the tool call, then the answer
assert_every_agent_ran(module, known.agents_ran | {"tool_calling"})
async def test_an_unknown_version_fails_the_tool_and_the_model_says_so_without_guessing(unknown):
returns = tool_returns(unknown)
assert returns and returns[0].outcome == "failed" # ToolFailed, not a retry
# However the model words it, the answer is about 3.99 and borrows nothing from the table.
answer = unknown.result.output.result.lower()
assert answer.strip() and "3.99" in answer
assert not [fact for fact in KNOWN_FACTS if fact in answer], answer
assert unknown.result.usage.requests >= 2 # the failed call, then the answer
async def test_the_demo_script_runs():
assert "result=" in await run_as_script("examples.tool_calling.agent")
Recorded run · gemini-3.1-flash-lite · 1 step · $0.0006
Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 996 tokens · $0.0006 · 2.3 s.
Model output varies between runs. Regenerate with uv run python scripts/record_example.py tool_calling.
Input
What changed in Python 3.13 compared with 3.12?
Steps
1. tool_calling
996 tokens · $0.0006
Prompt
What changed in Python 3.13 compared with 3.12?
What happened
- called python_release_notes({"version": "3.13"})
- called python_release_notes({"version": "3.12"})
- python_release_notes returned: Python 3.13 was released on October 7, 2024. Highlights: an experimental free-threaded build with the GIL disabled (PEP 703); an experimental JIT compiler (PEP 744); a new interactive interpreter with …
- python_release_notes returned: Python 3.12 was released on October 2, 2023. Highlights: cleaner generics syntax: type parameters and the type statement (PEP 695); f-strings can nest quotes and span lines (PEP 701); per-interpreter …
Output
{
"result": "Python 3.13, released on October 7, 2024, introduced several significant experimental features compared to Python 3.12 (released October 2, 2023).\n\nKey changes in Python 3.13 include:\n* **Free-threaded build:** An experimental mode with the Global Interpreter Lock (GIL) disabled (PEP 703).\n* **JIT Compiler:** An experimental Just-In-Time compiler (PEP 744).\n* **New Interactive Interpreter:** Features multi-line editing and color support.\n\nFor context, Python 3.12 had previously introduced improvements such as a cleaner syntax for generics (PEP 695), enhanced f-strings that allow nested quotes and multi-line expressions (PEP 701), and the ability to use per-interpreter GILs for subinterpreters (PEP 684).",
"versions": [
"3.12",
"3.13"
]
}
Result
run_tool_agent(...).output
{
"result": "Python 3.13, released on October 7, 2024, introduced several significant experimental features compared to Python 3.12 (released October 2, 2023).\n\nKey changes in Python 3.13 include:\n* **Free-threaded build:** An experimental mode with the Global Interpreter Lock (GIL) disabled (PEP 703).\n* **JIT Compiler:** An experimental Just-In-Time compiler (PEP 744).\n* **New Interactive Interpreter:** Features multi-line editing and color support.\n\nFor context, Python 3.12 had previously introduced improvements such as a cleaner syntax for generics (PEP 695), enhanced f-strings that allow nested quotes and multi-line expressions (PEP 701), and the ability to use per-interpreter GILs for subinterpreters (PEP 684).",
"versions": [
"3.12",
"3.13"
]
}