Skip to content

Code mode

Let the model write Python that calls your tools, in a sandbox.

In ordinary tool calling the model asks for one tool at a time and reads each result before asking for the next. That gets slow and expensive when a question needs dozens of lookups, and models add numbers badly in their heads. Code mode lets the model write a short program instead. Your tools are hidden behind a single run_code tool, the model writes Python that calls them in loops and does the arithmetic, and that code runs in a sandbox that can touch nothing except your tools. One or two model requests can do the work of dozens, and the totals are exact.

Use it when

  • A question needs many tool calls: looping over records, joining two lookups, aggregating.
  • It needs exact arithmetic, which models get wrong when they add numbers in their head.
  • You would rather have one or two model round trips than dozens.

Look elsewhere when

  • A question needs only one or two tool calls: tool_calling.
  • The model has to judge each result before deciding the next step: the code runs to the end without it.
  • The sandbox runs a subset of Python, with no file or network access, so keep the real logic in your tools.
plain tool calling   list ids → get_expense × 10 → get_exchange_rate × 3 → add it up by hand
code mode            run_code( a loop that fetches, converts and sums, exactly )

What it shows

  • The CodeMode capability (from pydantic-ai-harness) hides the agent's tools behind one run_code tool. The model writes Python that calls them as await get_expense(expense_id=...); each call really runs your tool on the host, with your deps. The tools are ordinary @agent.tool functions, so any agent's tools can be used this way
  • A real sandbox. The code runs in Monty, a minimal Python interpreter built for untrusted code, with no filesystem, environment, clock, network or subprocess. It is checked against your tools' signatures before it runs, and bounded by a time limit, a memory limit and a cap on tool calls per snippet (SANDBOX_LIMITS, MAX_TOOL_CALLS). A snippet that breaks a rule fails with an error the model reads and fixes; nothing happens on the host
  • Proved, not claimed. The offline tests run real hostile code in the real sandbox (reading and writing files, the environment, the clock, a socket, a subprocess, an infinite loop, a memory bomb, a runaway tool loop) and check the host: no file created, no secret returned, exactly the capped number of tool calls made, the loop stopped at its time limit
  • Exact answers, checked. The live tests compare the model's figures with totals worked out independently from the data (a total across five currencies; the employee with the most meal spend), and check deps.calls, a ledger of every tool call the code really made, to show the work was done
  • The payoff, measured: dozens of host tool calls from a model that needed two or three requests
  • A grounding check: an output validator rejects an answer whose subject isn't a real employee or expense id

See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.

uv run python scripts/add_agent.py code_mode --name expenses

This is the first example that needs a package of its own: add_agent.py installs pydantic-ai-harness[code-mode] with uv add. It is pre-1.0 and its minor version tracks Pydantic AI's (0.54 goes with 2.54), so upgrade the two together.

To adapt it, replace the expense data and the four tools with your own. Keep the tool signatures typed and documented, because the model's code is written against them.

Source

All of it is in examples/code_mode/.

"""Code mode: let the model write code that calls your tools, in a sandbox.

Use this pattern when:
- A question needs many tool calls (loop over records, join two lookups, aggregate)
- It needs exact arithmetic, which models get wrong when they add numbers in their head
- You want one or two model round trips instead of dozens

How it works: the `CodeMode` capability (from `pydantic-ai-harness`) hides the agent's tools behind a
single `run_code` tool. The model writes Python that calls them as `await get_expense(expense_id=...)`;
the code runs in **Monty**, a minimal Python interpreter built to run untrusted code. Each call the code
makes really runs your tool on the host, with your deps. Compare "how much did Maya spend on travel":

    plain tool calling   list ids → get_expense × 10 → get_exchange_rate × 3, then add them up by hand
    code mode            one `run_code` call: a loop that fetches, converts and sums, exactly

The sandbox has no filesystem, environment, clock, network or subprocess, and it is bounded: a time
limit, a memory limit, and a cap on how many tool calls one snippet may make. The code is also
type-checked against your tools' signatures before it runs, so a wrong argument is caught first. A
snippet that breaks a rule fails with an error the model reads and can fix; nothing happens on the
host. Those limits are set below (`SANDBOX_LIMITS`, `MAX_TOOL_CALLS`).

The expense data is invented and held in memory so the example runs anywhere. Replace the tools with
your own (an API, a database); `deps.calls` is a ledger of every tool call that really ran.
"""

from __future__ import annotations

from dataclasses import dataclass, field
from typing import Literal

from pydantic import BaseModel
from pydantic_ai import Agent, ModelRetry, RunContext, ToolFailed
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits
from pydantic_ai_harness import CodeMode

from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult

logger = get_logger(__name__)
LABEL = agent_label(__name__)  # names this agent's run spans in Logfire traces

# The model's requests, not the sandbox's tool calls: code mode needs few of them, so this is tight.
USAGE_LIMITS = UsageLimits(
    request_limit=8, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)

# --- The sandbox's limits ---
MAX_TOOL_CALLS = 60  # tool calls one `run_code` snippet may make (a runaway loop stops here)
SANDBOX_LIMITS = {
    "max_duration_secs": 2.0,  # CPU time per snippet; time spent waiting on tools does not count
    "max_memory": 100_000_000,  # bytes
}


# --- The data ---
@dataclass(frozen=True)
class Employee:
    id: str
    name: str


Category = Literal["travel", "meals", "lodging", "supplies"]


@dataclass(frozen=True)
class Expense:
    id: str
    employee_id: str
    category: Category  # spelled out, so the model's code sees the exact values to compare against
    amount: float
    currency: str  # a three-letter code such as "EUR"


EMPLOYEES: tuple[Employee, ...] = (
    Employee("E1", "Arjun"),
    Employee("E2", "Maya"),
    Employee("E3", "Tomas"),
)

# Replace with a real data source. Mixed currencies on purpose: the arithmetic is what code is for.
EXPENSES: tuple[Expense, ...] = (
    Expense("X101", "E1", "travel", 412.50, "USD"),
    Expense("X102", "E1", "meals", 38.20, "EUR"),
    Expense("X103", "E1", "travel", 129.00, "GBP"),
    Expense("X104", "E1", "meals", 22.75, "USD"),
    Expense("X105", "E1", "lodging", 640.00, "USD"),
    Expense("X106", "E1", "meals", 5400.00, "JPY"),
    Expense("X107", "E1", "supplies", 88.10, "CAD"),
    Expense("X108", "E1", "travel", 61.40, "EUR"),
    Expense("X201", "E2", "travel", 950.00, "EUR"),
    Expense("X202", "E2", "meals", 41.60, "USD"),
    Expense("X203", "E2", "travel", 18200.00, "JPY"),
    Expense("X204", "E2", "lodging", 420.00, "GBP"),
    Expense("X205", "E2", "travel", 75.25, "USD"),
    Expense("X206", "E2", "meals", 64.30, "CAD"),
    Expense("X207", "E2", "supplies", 29.99, "USD"),
    Expense("X208", "E2", "travel", 233.80, "GBP"),
    Expense("X209", "E2", "meals", 17.40, "EUR"),
    Expense("X301", "E3", "meals", 120.00, "USD"),
    Expense("X302", "E3", "travel", 305.40, "CAD"),
    Expense("X303", "E3", "meals", 87.50, "GBP"),
    Expense("X304", "E3", "lodging", 510.00, "EUR"),
    Expense("X305", "E3", "meals", 9800.00, "JPY"),
    Expense("X306", "E3", "supplies", 14.20, "USD"),
)

# How many US dollars one unit of each currency is worth.
RATES_TO_USD: dict[str, float] = {"USD": 1.0, "EUR": 1.08, "GBP": 1.27, "JPY": 0.0067, "CAD": 0.74}


# --- Dependencies ---
@dataclass
class ExpenseDeps:
    """Runtime dependencies for the expenses agent."""

    employees: tuple[Employee, ...] = EMPLOYEES
    expenses: tuple[Expense, ...] = EXPENSES
    rates: dict[str, float] = field(default_factory=lambda: dict(RATES_TO_USD))
    # Every tool call that really ran on the host, in order: ("get_expense", "X101"). The sandbox is
    # a boundary, so this is the ground truth for what the model's code actually did.
    calls: list[tuple[str, str]] = field(default_factory=list)


# --- Output type ---
class Answer(BaseModel):
    # `result` is the conventional output field in these examples; the generated
    # eval starter reads it when present (see evals/helpers.py).
    result: str
    amount_usd: float  # the dollar figure the question asked for
    subject: str  # the employee id or expense id the answer is about


# --- Agent ---
code_mode_agent: Agent[ExpenseDeps, Answer] = Agent(
    settings.model,
    name=LABEL,
    output_type=Answer,
    deps_type=ExpenseDeps,
    capabilities=[
        RaiseContentFilterError(),
        # Every tool below becomes a function the model's code can call; `run_code` is the only tool the
        # model sees. Pass `tools=[...]` to leave some as ordinary tool calls instead.
        CodeMode(max_tool_calls=MAX_TOOL_CALLS, resource_limits=SANDBOX_LIMITS),
    ],
    instructions=load_prompt("code_mode"),  # prompts/…; copied to agent/prompts/<name>.txt
)


# --- Tools: ordinary tools. Code mode makes them callable from the sandbox. ---
@code_mode_agent.tool
async def list_employees(ctx: RunContext[ExpenseDeps]) -> list[Employee]:
    """Every employee: their id and name."""
    ctx.deps.calls.append(("list_employees", ""))
    return list(ctx.deps.employees)


@code_mode_agent.tool
async def list_expense_ids(ctx: RunContext[ExpenseDeps], employee_id: str) -> list[str]:
    """The ids of one employee's expenses.

    Args:
        employee_id: The employee's id, such as "E2".

    Raises:
        ToolFailed: When there is no such employee.
    """
    ctx.deps.calls.append(("list_expense_ids", employee_id))
    if not any(e.id == employee_id for e in ctx.deps.employees):
        raise ToolFailed(f"There is no employee {employee_id!r}.")
    return [x.id for x in ctx.deps.expenses if x.employee_id == employee_id]


@code_mode_agent.tool
async def get_expense(ctx: RunContext[ExpenseDeps], expense_id: str) -> Expense:
    """One expense: its category, amount and currency.

    Args:
        expense_id: The expense id, such as "X101".

    Raises:
        ToolFailed: When there is no such expense.
    """
    ctx.deps.calls.append(("get_expense", expense_id))
    for expense in ctx.deps.expenses:
        if expense.id == expense_id:
            return expense
    raise ToolFailed(f"There is no expense {expense_id!r}.")


@code_mode_agent.tool
async def get_exchange_rate(ctx: RunContext[ExpenseDeps], currency: str) -> float:
    """How many US dollars one unit of `currency` is worth (multiply an amount by it).

    Args:
        currency: A three-letter code such as "EUR".

    Raises:
        ToolFailed: When the currency is not supported.
    """
    ctx.deps.calls.append(("get_exchange_rate", currency))
    try:
        return ctx.deps.rates[currency]
    except KeyError:
        raise ToolFailed(
            f"No rate for {currency!r}. Supported: {', '.join(sorted(ctx.deps.rates))}."
        ) from None


# --- Grounding ---
@code_mode_agent.output_validator
def check_the_subject_exists(ctx: RunContext[ExpenseDeps], answer: Answer) -> Answer:
    """The answer must point at a real employee or expense, so a made-up id is sent back."""
    known = {e.id for e in ctx.deps.employees} | {x.id for x in ctx.deps.expenses}
    if answer.subject not in known:
        raise ModelRetry(
            f"`subject` {answer.subject!r} is not an employee or expense id. Use an id from the data, "
            "such as an employee id like 'E2' or an expense id like 'X101'."
        )
    return answer


async def run_expenses(user_input: str, deps: ExpenseDeps | None = None) -> RunResult[Answer]:
    """Answer a question about the expense reports, using code to call the tools.

    Returns:
        A RunResult: `.output` is the `Answer`. The model's code is in `.all_messages()` (the
        `run_code` calls), and `deps.calls` records every tool call the code made.
    """
    if deps is None:
        deps = ExpenseDeps()
    logger.info("Running expenses agent", extra={"user_input": user_input})
    flow = Flow(USAGE_LIMITS)
    result = await flow.run(code_mode_agent, user_input, deps=deps)
    return flow.finish(result.output)


if __name__ == "__main__":
    import asyncio

    configure_logging()
    deps = ExpenseDeps()
    result = asyncio.run(
        run_expenses("What is the total of Maya's travel expenses, in US dollars?", deps)
    )
    print(
        result.output, f"| {len(deps.calls)} tool calls in {result.usage.requests} model requests"
    )
You answer questions about employee expense reports, using the tools provided.

- Work with code: call `run_code` and, inside it, call the tools in loops and do the arithmetic in
  Python. Never add numbers yourself and never guess; the code gives exact answers.
- Expenses are in different currencies. Convert every amount to US dollars with
  `get_exchange_rate` (the rate multiplies the amount to give dollars) before totalling.
- The sandbox runs a subset of Python. Use plain `for` loops and dictionaries; avoid `next()` over a
  generator expression and other advanced features. Tool results come back as dicts.
- Round money to two decimal places.
- Fill `amount_usd` with the dollar figure the question asks for, and `subject` with the id of the
  employee or expense the answer is about (for example "E2").
title = "Code mode"
pattern = "code_mode"
summary = "Let the model write Python that calls your tools in a sandbox (Monty): many tool calls and exact arithmetic in one or two model requests, with hard limits on what the code can do."
smoke_input = "What is the total of Maya's travel expenses, in US dollars?"
# A real model must work through code (the sandbox's one tool) for this (checked by the release check).
expected_tools = ["run_code"]
# Code mode comes from pydantic-ai-harness, which tracks Pydantic AI's version (0.54 goes with 2.54).
dependencies = ["pydantic-ai-harness[code-mode]>=0.54,<1"]

# TestModel's generated answer cites a subject that is not a real id, which the output validator
# rejects, so the smoke tests supply one that is.
[smoke.code_mode_agent]
output = { result = "Maya spent a total of 0 dollars.", amount_usd = 0.0, subject = "E2" }

[entrypoint]
deps = "ExpenseDeps"
run = "run_expenses"
"""The tools, the answer check, and — the point of the example — what the sandbox lets code do.

Monty is real and local, so these run actual model-style code in the actual sandbox with a scripted
model, and check what happened on the *host* (the tool-call ledger, files, time). Needs
pydantic-ai-harness (declared in example.toml), so it is skipped without it and run in the example's own
environment by scripts/release_check.py.
"""

import time
from pathlib import Path

import pytest

pytest.importorskip("pydantic_ai_harness", reason="needs pydantic-ai-harness[code-mode]")

from pydantic_ai import ModelRetry, RunContext, ToolFailed  # noqa: E402
from pydantic_ai.messages import (  # noqa: E402
    ModelResponse,
    RetryPromptPart,
    ToolCallPart,
    ToolReturnPart,
)
from pydantic_ai.models.function import AgentInfo, FunctionModel  # noqa: E402
from pydantic_ai.models.test import TestModel  # noqa: E402
from pydantic_ai.usage import RunUsage  # noqa: E402

from examples.code_mode.agent import (  # noqa: E402
    EMPLOYEES,
    EXPENSES,
    MAX_TOOL_CALLS,
    RATES_TO_USD,
    SANDBOX_LIMITS,
    Answer,
    ExpenseDeps,
    check_the_subject_exists,
    code_mode_agent,
    get_exchange_rate,
    get_expense,
    list_employees,
    list_expense_ids,
    run_expenses,
)


def ctx(deps: ExpenseDeps | None = None) -> RunContext[ExpenseDeps]:
    return RunContext(deps=deps or ExpenseDeps(), model=TestModel(), usage=RunUsage())


def truth(employee_id: str, category: str) -> float:
    """The right answer, worked out independently of the model and of the sandbox."""
    return round(
        sum(
            x.amount * RATES_TO_USD[x.currency]
            for x in EXPENSES
            if x.employee_id == employee_id and x.category == category
        ),
        2,
    )


# --- The data ---


def test_the_data_is_consistent_so_the_ground_truth_can_be_trusted():
    assert len({x.id for x in EXPENSES}) == len(EXPENSES)  # ids are unique
    assert {x.employee_id for x in EXPENSES} <= {e.id for e in EMPLOYEES}
    assert {x.currency for x in EXPENSES} <= set(RATES_TO_USD)
    assert truth("E2", "travel") == 1520.12 and max(
        (truth(e.id, "meals"), e.id) for e in EMPLOYEES
    ) == (296.78, "E3")


# --- The tools, called directly; each records itself in the ledger ---


async def test_employees_are_listed_and_recorded():
    deps = ExpenseDeps()
    assert [e.name for e in await list_employees(ctx(deps))] == ["Arjun", "Maya", "Tomas"]
    assert deps.calls == [("list_employees", "")]


async def test_an_employees_expense_ids_are_listed():
    deps = ExpenseDeps()
    ids = await list_expense_ids(ctx(deps), "E2")
    assert ids == [x.id for x in EXPENSES if x.employee_id == "E2"] and len(ids) == 9
    assert deps.calls == [("list_expense_ids", "E2")]


async def test_an_unknown_employee_is_a_terminal_failure():
    with pytest.raises(ToolFailed, match="no employee 'E9'"):
        await list_expense_ids(ctx(), "E9")


async def test_one_expense_is_returned_whole():
    expense = await get_expense(ctx(), "X201")
    assert (expense.employee_id, expense.category, expense.amount, expense.currency) == (
        "E2",
        "travel",
        950.0,
        "EUR",
    )


async def test_an_unknown_expense_is_a_terminal_failure():
    with pytest.raises(ToolFailed, match="no expense 'X999'"):
        await get_expense(ctx(), "X999")


async def test_a_rate_is_looked_up_and_an_unknown_currency_lists_the_supported_ones():
    assert await get_exchange_rate(ctx(), "EUR") == 1.08
    with pytest.raises(ToolFailed, match="Supported: CAD, EUR, GBP, JPY, USD"):
        await get_exchange_rate(ctx(), "XYZ")


def test_each_run_gets_its_own_ledger():
    assert ExpenseDeps().calls is not ExpenseDeps().calls


# --- The answer check ---


@pytest.mark.parametrize("subject", ["E2", "X101"])
def test_an_employee_or_expense_id_is_accepted(subject):
    answer = Answer(result="r", amount_usd=1.0, subject=subject)
    assert check_the_subject_exists(ctx(), answer) is answer


def test_an_id_that_does_not_exist_is_sent_back():
    with pytest.raises(ModelRetry, match="'Maya' is not an employee or expense id"):
        check_the_subject_exists(ctx(), Answer(result="r", amount_usd=1.0, subject="Maya"))


# --- A scripted model that writes code ---


def scripted(*steps: dict, seen: list | None = None):
    """Follow `steps`: {"code": ...} calls `run_code`; {"output": {...}} gives the final answer."""
    remaining = list(steps)

    def model_fn(messages, info: AgentInfo) -> ModelResponse:
        if seen is not None:
            seen.append((messages, info))
        step = remaining.pop(0)
        if "code" in step:
            return ModelResponse(parts=[ToolCallPart("run_code", {"code": step["code"]})])
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, step["output"])])

    return FunctionModel(model_fn)


def done(amount: float = 0.0, subject: str = "E2") -> dict:
    return {"output": {"result": "done", "amount_usd": amount, "subject": subject}}


def run_code_results(result) -> list[str]:
    return [
        str(part.content)
        for message in result.all_messages()
        for part in message.parts
        if isinstance(part, ToolReturnPart) and part.tool_name == "run_code"
    ]


def retry_messages(result) -> list[str]:
    return [
        str(part.content)
        for message in result.all_messages()
        for part in message.parts
        if isinstance(part, RetryPromptPart)
    ]


MAYA_TRAVEL = """
employees = await list_employees()
maya = ''
for e in employees:
    if e['name'] == 'Maya':
        maya = e['id']
total = 0.0
rates = {}
for expense_id in await list_expense_ids(employee_id=maya):
    expense = await get_expense(expense_id=expense_id)
    if expense['category'] == 'travel':
        if expense['currency'] not in rates:
            rates[expense['currency']] = await get_exchange_rate(currency=expense['currency'])
        total += expense['amount'] * rates[expense['currency']]
round(total, 2)
"""


async def test_the_model_sees_one_tool_and_it_is_run_code():
    seen: list = []
    with code_mode_agent.override(model=scripted(done(), seen=seen)):
        await run_expenses("anything")
    tools = seen[0][1].function_tools
    assert [t.name for t in tools] == ["run_code"]  # the five real tools are not offered directly
    description = tools[0].description or ""
    for name in ("list_employees", "list_expense_ids", "get_expense", "get_exchange_rate"):
        assert name in description  # …but the model is told it can call them from its code


async def test_one_snippet_does_the_work_of_dozens_of_tool_calls_and_gets_the_exact_answer():
    deps = ExpenseDeps()
    with code_mode_agent.override(model=scripted({"code": MAYA_TRAVEL}, done(1520.12))):
        result = await run_expenses("Maya's travel total?", deps)

    assert run_code_results(result) == ["1520.12"] and float(run_code_results(result)[0]) == truth(
        "E2", "travel"
    )
    assert result.usage.requests == 2  # the snippet, then the answer
    # The ledger is what really ran on the host: all nine of Maya's expenses, fetched by the code.
    fetched = [arg for name, arg in deps.calls if name == "get_expense"]
    assert fetched == [x.id for x in EXPENSES if x.employee_id == "E2"]
    assert len(deps.calls) >= 12  # many tool calls from a single model request
    assert result.output.amount_usd == 1520.12


async def test_independent_lookups_can_run_concurrently_from_one_snippet():
    code = """
import asyncio
rates = await asyncio.gather(*[get_exchange_rate(currency=c) for c in ['EUR', 'GBP', 'JPY']])
rates
"""
    deps = ExpenseDeps()
    with code_mode_agent.override(model=scripted({"code": code}, done())):
        result = await run_expenses("rates", deps)
    assert run_code_results(result) == ["[1.08, 1.27, 0.0067]"]
    assert sorted(arg for _, arg in deps.calls) == ["EUR", "GBP", "JPY"]


async def test_state_survives_between_snippets_like_a_repl():
    steps = [{"code": "n = len(await list_employees())"}, {"code": "n * 10"}, done()]
    with code_mode_agent.override(model=scripted(*steps)):
        result = await run_expenses("state")
    assert run_code_results(result)[-1] == "30"


async def test_a_tool_that_fails_inside_the_code_tells_the_model_why():
    code = "await get_expense(expense_id='X999')"
    with code_mode_agent.override(model=scripted({"code": code}, done())):
        result = await run_expenses("missing")
    assert any("no expense 'X999'" in m for m in retry_messages(result))


async def test_the_model_recovers_from_a_mistake_in_its_own_code():
    steps = [
        {"code": "expense = await get_expense(expense_id='X101')\nexpense['amount_in_dollars']"},
        {"code": "expense = await get_expense(expense_id='X101')\nexpense['amount']"},
        done(412.5, "X101"),
    ]
    deps = ExpenseDeps()
    with code_mode_agent.override(model=scripted(*steps)):
        result = await run_expenses("one expense", deps)

    assert len(retry_messages(result)) == 1 and "amount_in_dollars" in retry_messages(result)[0]
    assert run_code_results(result) == ["412.5"] and result.output.subject == "X101"


# --- What the sandbox does not allow: each hostile snippet is refused, and the host is untouched ---


async def hostile(code: str, deps: ExpenseDeps | None = None):
    """Run one snippet of untrusted code; the model then recovers and answers."""
    deps = deps or ExpenseDeps()
    with code_mode_agent.override(model=scripted({"code": code}, done())):
        result = await run_expenses("do something unsafe", deps)
    return result, deps


async def test_code_cannot_read_a_file_on_the_host():
    result, deps = await hostile("open('/etc/passwd').read()")
    # Refused before it runs: the sandbox has no `open` at all (the exact wording is Monty's).
    assert retry_messages(result) and not run_code_results(result)
    assert "root:" not in " ".join(retry_messages(result))  # and nothing from the file came back
    assert deps.calls == []


async def test_pathlib_can_be_imported_but_cannot_touch_the_filesystem():
    result, _ = await hostile("from pathlib import Path\nPath('/etc/passwd').read_text()")
    assert any("PermissionError" in m for m in retry_messages(result))
    assert not run_code_results(result)


async def test_code_cannot_write_a_file_on_the_host(tmp_path: Path):
    target = tmp_path / "pwned.txt"
    result, _ = await hostile(f"open({str(target)!r}, 'w').write('x')")
    assert retry_messages(result) and not run_code_results(result)
    assert not target.exists()  # the file was never created on the real filesystem


async def test_pathlib_cannot_write_a_file_either(tmp_path: Path):
    target = tmp_path / "pwned.txt"
    result, _ = await hostile(f"from pathlib import Path\nPath({str(target)!r}).write_text('x')")
    assert any("PermissionError" in m for m in retry_messages(result))
    assert not target.exists()


async def test_code_cannot_list_a_directory():
    result, _ = await hostile("import os\nos.listdir('/')")
    assert any("PermissionError" in m for m in retry_messages(result))


async def test_code_cannot_read_environment_variables(monkeypatch):
    monkeypatch.setenv("SECRET_TOKEN", "hunter2")
    result, _ = await hostile("import os\nos.environ['SECRET_TOKEN']")
    assert retry_messages(result)
    assert "hunter2" not in " ".join(retry_messages(result) + run_code_results(result))


async def test_code_cannot_read_the_clock():
    result, _ = await hostile("import time\ntime.time()")
    assert any("not supported in this environment" in m for m in retry_messages(result))


@pytest.mark.parametrize("module", ["socket", "subprocess"])
async def test_code_cannot_reach_the_network_or_start_a_process(module):
    result, _ = await hostile(f"import {module}\n{module}")
    assert any(f"Cannot resolve imported module `{module}`" in m for m in retry_messages(result))


async def test_a_runaway_loop_is_stopped_at_the_time_limit():
    started = time.perf_counter()
    result, _ = await hostile("while True:\n    pass")
    elapsed = time.perf_counter() - started
    assert any("max_duration_secs" in m for m in retry_messages(result))
    assert SANDBOX_LIMITS["max_duration_secs"] <= elapsed < SANDBOX_LIMITS["max_duration_secs"] + 3


async def test_a_memory_bomb_is_refused():
    result, _ = await hostile("blob = 'x' * (10**10)\nlen(blob)")
    assert any("MemoryError" in m for m in retry_messages(result))


async def test_a_runaway_tool_loop_is_capped_and_only_that_many_calls_reach_the_host():
    code = "for i in range(1000):\n    await get_exchange_rate(currency='EUR')"
    result, deps = await hostile(code)
    assert any(f"allows {MAX_TOOL_CALLS} nested tool calls" in m for m in retry_messages(result))
    assert len(deps.calls) == MAX_TOOL_CALLS  # exactly the cap, not 1000


async def test_a_refused_snippet_does_not_end_the_run_the_model_can_carry_on():
    result, _ = await hostile("open('/etc/passwd')")
    assert result.output.result == "done"  # it recovered and answered


# --- The configuration itself ---


def test_the_sandbox_limits_are_finite_and_set():
    assert 0 < SANDBOX_LIMITS["max_duration_secs"] <= 10
    assert 0 < SANDBOX_LIMITS["max_memory"] <= 1_000_000_000
    assert 0 < MAX_TOOL_CALLS <= 200
"""Live check: the model solves real questions by writing code, and the numbers are exactly right.

The tool-call ledger (`deps.calls`) is the ground truth for what the model's code did on the host,
and the answers are compared with figures worked out independently from the data. Needs
pydantic-ai-harness (declared in example.toml); run by scripts/release_check.py in the example's own
environment. Run with `pytest -m eval`.
"""

import pytest

pytest.importorskip("pydantic_ai_harness", reason="needs pydantic-ai-harness[code-mode]")

from evals.trace import traced_run  # noqa: E402
from examples.code_mode import agent as module  # noqa: E402
from examples.live_support import assert_every_agent_ran, run_as_script  # noqa: E402

pytestmark = pytest.mark.eval

CENT = 0.01


def worked_out(employee_id: str, category: str) -> float:
    """The right answer, from the data, without the model or the sandbox."""
    return round(
        sum(
            x.amount * module.RATES_TO_USD[x.currency]
            for x in module.EXPENSES
            if x.employee_id == employee_id and x.category == category
        ),
        2,
    )


async def ask(question: str):
    """Run the agent with a fresh ledger: returns the traced run and the deps it used."""
    deps = module.ExpenseDeps()

    async def helper(text: str):
        return await module.run_expenses(text, deps)

    return await traced_run(helper, question), deps


def fetched(deps: module.ExpenseDeps) -> set[str]:
    """The expense ids the model's code actually looked up on the host."""
    return {arg for name, arg in deps.calls if name == "get_expense"}


@pytest.fixture(scope="module")
async def travel():
    return await ask("What is the total of Maya's travel expenses, in US dollars?")


@pytest.fixture(scope="module")
async def meals():
    return await ask(
        "Which employee spent the most on meals, in US dollars? Give their id and their meal total."
    )


@pytest.fixture(scope="module")
async def one_expense():
    return await ask("What is expense X203 worth in US dollars?")


async def test_a_total_across_currencies_is_exactly_right_and_came_from_code(travel):
    traced, _ = travel
    assert "run_code" in traced.tools_called  # the sandbox's one tool: the work was done in code
    output = traced.result.output
    assert output.amount_usd == pytest.approx(worked_out("E2", "travel"), abs=CENT)  # 1520.12
    assert output.subject == "E2"


async def test_the_code_really_fetched_every_record_it_needed(travel):
    _, deps = travel
    mayas_travel = {
        x.id for x in module.EXPENSES if x.employee_id == "E2" and x.category == "travel"
    }
    assert mayas_travel <= fetched(deps)  # no expense was skipped or guessed


async def test_comparing_every_employee_picks_the_right_one(meals):
    traced, deps = meals
    output = traced.result.output
    assert (
        output.subject == "E3"
    )  # not the employee with the most meals, but the largest in dollars
    assert output.amount_usd == pytest.approx(worked_out("E3", "meals"), abs=CENT)  # 296.78
    every_meal = {x.id for x in module.EXPENSES if x.category == "meals"}
    assert every_meal <= fetched(deps)  # all three employees' meals, converted correctly


async def test_a_single_conversion_is_exact(one_expense):
    traced, _ = one_expense
    expected = round(18200 * module.RATES_TO_USD["JPY"], 2)  # 121.94
    assert traced.result.output.amount_usd == pytest.approx(expected, abs=CENT)
    assert traced.result.output.subject == "X203"


async def test_one_or_two_model_requests_do_the_work_of_dozens_of_tool_calls(travel, meals):
    """The reason for the pattern. Plain tool calling needs a model round trip per step."""
    for traced, deps in (travel, meals):
        requests = traced.result.usage.requests
        assert len(deps.calls) >= 10
        assert len(deps.calls) >= 3 * requests, f"{len(deps.calls)} calls in {requests} requests"


async def test_the_sandbox_was_the_only_way_in(travel, meals, one_expense):
    """The model was offered `run_code` and nothing else, so every host call came from its code."""
    for traced, deps in (travel, meals, one_expense):
        assert traced.tools_called.count("run_code") >= 1
        assert deps.calls  # and the code did reach the real tools


async def test_the_agent_ran(travel, meals, one_expense):
    ran = travel[0].agents_ran | meals[0].agents_ran | one_expense[0].agents_ran
    assert_every_agent_ran(module, ran)


async def test_the_demo_script_runs():
    out = await run_as_script("examples.code_mode.agent")
    assert "amount_usd=" in out and "tool calls in" in out

Recorded run · gemini-3.1-flash-lite · 1 step · $0.0012

Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 3,083 tokens · $0.0012 · 2.5 s. Model output varies between runs. Regenerate with uv run python scripts/record_example.py code_mode.

Input

What is the total of Maya's travel expenses, in US dollars?

Steps

1. code_mode 3,083 tokens · $0.0012

Prompt

What is the total of Maya's travel expenses, in US dollars?

What happened - ran this code in the sandbox:

```python
async def get_maya_travel_total():
    employees = await list_employees()
    maya_id = None
    for emp in employees:
        if emp['name'] == 'Maya':
            maya_id = emp['id']
            break

    if not maya_id:
        return "Maya not found"

    total_usd = 0.0
    expense_ids = await list_expense_ids(employee_id=maya_id)

    for exp_id in expense_ids:
        expense = await get_expense(expense_id=exp_id)
        if expense['category'] == 'travel':
            amount = expense['amount']
            currency = expense['currency']
            if currency == 'USD':
                rate = 1.0
            else:
                rate = await get_exchange_rate(currency=currency)
            total_usd += amount * rate

    return {"total": round(total_usd, 2), "id": maya_id}

await get_maya_travel_total()
```
  • run_code returned: {"total": 1520.12, "id": "E2"}

Output

{
  "result": "The total of Maya's travel expenses is $1520.12.",
  "amount_usd": 1520.12,
  "subject": "E2"
}

Result

run_expenses(...).output

{
  "result": "The total of Maya's travel expenses is $1520.12.",
  "amount_usd": 1520.12,
  "subject": "E2"
}