Skip to content

Design notes

This repo is a template: a clean starting point for a production-quality Pydantic AI agent, meant to be cloned and reshaped into a specific agent. The code is the source of truth.

Commands

uv sync --group dev            # install deps
uv run pytest                  # unit tests only (TestModel, no API key needed)
uv run pytest -m eval          # all evals — pass/fail AND LLM-judge (needs a real API key, costs money)
uv run ruff check .            # lint
uv run ruff format .           # format

There is no separate llm_judge marker. Everything in evals/ carries only @pytest.mark.eval, so -m eval runs it all in one shot — there's no cheaper eval-only subset to reach for.

asyncio_mode = "auto" is set in pyproject.toml, so async tests need no @pytest.mark.asyncio decorator — don't add them back. The loop scopes are also pinned to session there: agents are module-level, so each provider HTTP client binds to the loop of the first test that uses it, and per-test loops (the default) make every later real model call — i.e. every eval after the first — fail with Event loop is closed. tests/test_event_loop.py guards this.

Changelog

CHANGELOG.md follows Keep a Changelog and is written for people who cloned the template. Any user-visible change (new or changed setting, default, convention, scaffold output, or anything a cloner must act on) adds a line under ## [Unreleased] in the same commit, grouped as Added / Changed / Fixed / Removed, with an "Upgrade notes" entry when action is required. To release: bump version in pyproject.toml (and uv lock), rename [Unreleased] to the new version with its date, add a fresh empty [Unreleased], update the compare links at the bottom, and tag vX.Y.Z.

Releasing

Before tagging, run uv run python scripts/release_check.py (Docker must be running if any example declares services) (makes real model calls: needs the provider key for AGENT_MODEL and costs money — pennies per example on a small model). It exists because TestModel can't tell you an example works: it ignores what tools return and what instructions say, so a placeholder tool or a half-written prompt passes every offline test (the tool_calling example once shipped a tool that echoed its query, and a real model looped on it until the request limit). The gate:

  1. runs the offline suite under coverage;
  2. runs each example in its own uv run environment (only its declared dependencies layered on), under a hard spend cap from its cost_budget_usd: first its live tests (examples/<name>/test_live.py, marked eval — real-model assertions on behavior, plus a requirement that every agent the example defines actually ran, read from spans via evals/trace.py so workers inside tool calls count, plus a run of the module's if __name__ == "__main__" demo), then a smoke run (scripts/record_example.py: the run returns a RunResult, steps are labeled for the example, the manifest's expected_tools were really called, cost stayed in budget);
  3. fails unless every line of every examples/*/agent.py was executed by the offline and live tests together (coverage report, fail_under = 100 in pyproject.toml) — so no example code is untested;
  4. prints pass/fail/unverified counts, tokens and spend, and exits non-zero unless everything passed.
  5. when the whole library passed (all examples, tests included, nothing unverified), writes badges/coverage.json, which the README's shields.io endpoint badge reads; it records the coverage percentage (the badge reads coverage: 100%) and is never written by a partial run (badge_is_earned in scripts/release_check.py). Commit it with the release.

The gate checks whatever model AGENT_MODEL names, so you can reconfigure .env freely: the live tests assert behavior that holds for any capable model (tool outcomes, structure, which agents ran, no invented data), not one model's wording or request counts, and nothing in the gate names a provider. What a model must be able to do is the floor: call tools and return valid structured output. A model that can't (small local ones often return invalid JSON) fails the gate with Pydantic validation errors, which is a verdict on the model, not the example. The model that recorded each transcript is in its header, and --record overwrites them, so record with the model you want the docs to show.

--record also rewrites each example's sample_run.md (a transcript built from the RunResult; a failed run never overwrites the last good one — commit the result). scripts/record_example.py <name> records one example. Run a single example's live tests with uv run pytest -m eval examples/<name>. When you add an example, give it a test_live.py (copy the closest existing one): inputs with an unambiguous expected result, semantic assertions, every agent exercised, and the demo script. The gate is deliberately a manual, local step — it spends money and needs your key — so it is not wired into CI (.github/workflows/ci.yml runs only the offline suite). Run it yourself, with the provider key in .env or the environment; the model comes from AGENT_MODEL (the committed transcripts were recorded with google:gemini-3.1-flash-lite). An example that declares services is started with docker compose up --build --wait for the duration of its checks and always torn down; if Docker isn't usable it is reported as unverified with the reason (and fails the check unless --allow-unverified), never as passed.

Documentation site

The docs site (MkDocs Material, published to GitHub Pages by .github/workflows/docs.yml, served at https://tmtabor.io/agent-template/) is generated by docs/gen_pages.py from the repo's own files, with two hand-written exceptions. The sources: the README (split into pages by its ## sections; the home page is its header, pitch and sales sections), AGENTS.md (the "Design notes" page), CHANGELOG.md, CONTRIBUTING.md, MAINTAINING.md, each example (its README, sample_run.md, and its source in tabs), and the two guides in docs/pages/ ("Which pattern should I use?" and the FAQ), which are the only hand-written pages. Links are written repo-relative, like everything else, and rewritten to site pages or GitHub (including README.md#some-section, which lands on the page that now holds the section). So to change the docs, change those files. Consequences worth knowing: (1) the README's ## section titles are an interface: the ones in README_PAGES and HOME_SECTIONS in docs/gen_pages.py must exist, and renaming one makes the build fail loudly (SectionNotFound) until the generator is updated; (2) the README's "What are you building?" section is the home page's tabs: each ### under it becomes a tab on the site and an ordinary headed section on GitHub, so keep each one self-contained; (3) the README opens with a centred HTML header (logo in <picture>, tagline, badges, links): the site swaps the <picture> for Material's #only-light / #only-dark images, drops the links back to the site itself, and turns raw-HTML .md links into the directory URLs MkDocs serves, because MkDocs only fixes Markdown links; (4) the logo SVGs and the favicon are in docs/assets/ (the wordmark is drawn as outlines, so it needs no font); (5) a new example gets its page and nav entry automatically, needs a sample_run.md (tests/test_examples.py enforces it) and optionally a slot in DISPLAY_ORDER in scripts/example_manifest.py; (6) examples/README.md is generated (uv run python scripts/examples_index.py) and tests/test_docs.py fails if it is stale; (7) docs/pages/ is not published as it is, and on a pattern page the recorded run is a closed collapsible block below the source, titled with the run's model, steps and cost, behind an anchor that docs/assets/extra.js opens when a link points at #recorded-run (the README's "See it run" line); (8) write for GitHub but expect MkDocs: the generator adds the blank line MkDocs needs before a list that follows a **bold** line, but anything else GitHub forgives and MkDocs doesn't should be checked by looking at the page (mkdocs serve), because --strict only catches broken links and a raw-HTML link is not checked at all. The site builds with mkdocs build --strict, so a broken link fails the build, and tests/test_docs.py checks every internal link and nav entry without needing MkDocs. mkdocs is pinned <2: MkDocs 2.0 will break plugins and themes, and properdocs (a continuation of 1.x that Material already pulls in) is the drop-in if you need to leave 1.x. The docs dependency group is separate from dev; add_agent.py --prune removes the docs, their tests, that group, and CONTRIBUTING.md and MAINTAINING.md. The how-to for maintainers is in MAINTAINING.md.

Writing a new example

Everything a pattern example must have, because the gate (see Releasing) checks all of it and a TestModel can hide a broken example from every offline test:

  1. agent.py with a run_* helper returning a RunResult, a deps dataclass constructible with no arguments, and name=LABEL on every Agent (f"{LABEL}.role" for helpers) — keep every agent reachable from module scope. No placeholder tools, instructions or data: a real model will loop on a tool that echoes its input. Make the domain invented or deterministic, so a correct answer proves the pattern worked (RAG's fictional store policies, MCP's date arithmetic).
  2. example.toml: smoke_input that makes a real model exercise the pattern, expected_tools, and a [smoke.<agent>] table for any agent whose validators reject TestModel's junk (output = {...}), whose bool defaults the flow needs flipped, or that must call tools. A dependencies entry for anything the root pyproject.toml lacks.
  3. test_example.py: offline tests with scripted FunctionModels covering every branch — the gate requires 100% line coverage of agent.py over offline and live tests together. Test the side effect, not the model's words (the refund ledger, the retrieved ids), and any ground truth the example defines.
  4. test_live.py (marked eval): inputs with an unambiguous expected result, assertions that hold for any capable model (no exact request counts, no one model's phrasing), a requirement that every agent ran (assert_every_agent_ran), and a run of the __main__ demo. When the example has a pause, a block or a rejection, observe it in the spans or the ground truth, not just the final text.
  5. README.md (with the **See it run:** line), added to DISPLAY_ORDER in scripts/example_manifest.py, then uv run python scripts/examples_index.py, then uv run python scripts/release_check.py <name> --record.

An example with dependencies or test_dependencies is skipped by the generic tests in the default environment (an ImportError from a declared dependency is a skip, not a failure) and runs, with the generic and add_agent tests, in its own uv run --with environment inside the gate. dependencies are what using the agent needs (add_agent.py installs them); test_dependencies are what only its tests need (for instance the server half of a library, to run a service locally). If a library raises a plain ImportError when an extra is missing (as fastmcp does for its server half), probe for that half in the test file rather than the package name.

An example that runs untrusted code (code_mode is the model) must test the sandbox with real hostile input, not only the happy path, and check the host, not the sandbox's own error message: that no file appeared, no secret came back, no tool ran beyond its cap, and a runaway loop stopped at its time limit. Assert on the invariant rather than Monty's wording (it has no open at all until os is imported, then raises PermissionError). Monty is a Python subset, and it type-checks the model's code against the tool signatures before running it: for example next() over a generator expression fails, and x = None later passed as a str is rejected. Put what the model needs to avoid those in the prompt or the tool types (Literal categories stopped it comparing against 'Travel'), and measure the effect in model requests. pydantic-ai-harness is pre-1.0 and its minor version tracks Pydantic AI's (0.54 goes with 2.54), so bump them together; the dependencies pin is >=0.54,<1.

A planner-executor example (planner_executor is the model) must enforce the plan's rules in code and test each rule with a plan that breaks it, because a real model does not reliably follow them: Gemini wrote "compare them" steps with depends_on left empty in 9 of 10 plans, even with the rule spelled out and an example in the prompt, and the answers still came out right (the executors had the lookup tool and the synthesizer did the arithmetic), so a correct answer proves nothing about the plan. Check the plan's shape in the live tests (independent steps in one wave, a step that depends on others), not only the answer. The text a rejected plan is sent back with is part of the prompt: saying what was wrong made the model repeat the same plan three times; adding what to change (computed by the code) fixed it on the first retry in 11 of 12 runs. Prove the scheduling with scripted models that record what each executor was told and how many ran at once, and mutate the implementation (run steps one at a time, leak every result to every executor, stop the skip cascade early) to see that the tests notice.

A durable example (temporal is the model) must be tested against the real Temporal server, because Temporal's own test server downloads a binary at run time and a hermetic suite cannot: its tests skip without the service's address and the release check supplies it. Assert on the server's event history (fetch_history_events), not on the model's words: the model-request activity must show one start per request, and a retried tool shows only its final attempt number (failed attempts of a retried activity are not events), so the failures themselves belong in a ledger your own tool keeps. Prove crash recovery with a real SIGKILL of a worker in a separate process as well as a tidy shutdown (the server notices a dead worker only when the activity's timeout passes). Things that bite: a workflow cannot be defined in __main__, so the demo block re-imports the module under its real name; the agent's name (and therefore its activity names) must be the same in the worker and in the workflow's sandbox; the template's own modules must be imported inside workflow.unsafe.imports_passed_through() because the sandbox forbids what they do at import time; two workflows cannot share an agent; a bug in workflow code makes Temporal retry forever, so every run needs an execution timeout; and Pydantic AI's invoke_agent span does not appear for a run inside a workflow, so use the history as the evidence that the agent ran.

An example that needs a running service (mcp_tools is the model; temporal is the image-only variant) ships service/ (the server and a Dockerfile, or just a compose file naming a pinned published image, never latest; a docker-compose.yml that publishes the port as "127.0.0.1::<port>" and has a healthcheck, so up --wait returns when it is ready and nothing is exposed beyond this machine) and declares it in example.toml (services, and a [service.<name>] table with the container port, the env variable that receives its address, and a url template). The agent reads the address from deps, so tests and the gate can point it anywhere. Three layers of testing, none needing Docker until the last: the server's functions directly; the server run as a local subprocess on a free port for the offline tests (a fixture in the example's conftest.py); and the real container, which only the release check starts. Generic tests that run the agent skip such an example unless its address variable is set (import_example(example, running=True)), so the default suite never needs Docker.

Making it yours

The intended customization sequence, roughly in order:

  1. Add an agent: uv run python scripts/add_agent.py (menu) or add_agent.py <example> --name <name>. The template ships with no agents. The script copies an example from examples/ into agent/agents/<name>.py, its prompt into agent/prompts/, and scaffolds tests/test_agents_<name>.py and evals/test_<name>.py. Run it once per agent; each can use a different pattern. Import agents directly from their own modules — there is no shared "primary" agent.
  2. Edit the prompt: agent/prompts/<name>.txt, loaded via load_prompt("<name>"). Add more .txt files beside it and load them the same way.
  3. Define the output schema: replace the placeholder fields on the output model in your agent module. The examples keep a result: str field, which the eval starters and the web-UI skill read when present (evals/helpers.py:output_text falls back to str(output)). Keep the schema as flat as the data actually requires — don't wrap a single field in its own object (e.g. prefer list[str] over list[{text: str}]) unless a second field genuinely needs to travel with it. This matters more the smaller/weaker settings.model is: a schema a frontier model satisfies without issue can reliably burn the retry budget on a small local model (e.g. an ollama: model) if it adds nesting the data doesn't need. If output validation keeps failing against the configured model, check whether the schema is more nested than necessary before assuming it's a prompting problem.
  4. Add tools: copy the pattern in agent/tools/example.py, register with @<your_agent>.tool.
  5. Tune USAGE_LIMITS in your agent module: request_limit caps model round-trips per run, total_tokens_limit caps tokens. An optional USD spend cap comes from AGENT_COST_LIMIT (off by default). The supervisor shares its budget with workers via usage=ctx.usage.
  6. Grow the evals: add cases to evals/fixtures/<name>.json (picked up by the dataset eval in evals/test_<name>.py automatically) and adapt the judge criteria there. Fixtures also accept optional expected_tools, expected_trajectory and expected_arguments keys, which become span-based behavioral evaluators (ToolCorrectness, TrajectoryMatch, ArgumentCorrectness); every case gets MaxModelRequests/MaxToolCalls budgets. The shared evaluators and runner are in evals/helpers.py.
  7. Drop what you don't need: uv run python scripts/add_agent.py --prune deletes the examples you didn't use (keeping blank), the docs and their tests. Nothing under agent/ or evals/ depends on examples/.

Non-obvious architecture

  • agent/config.py validates at import time, not at call time. Settings has a model_validator that calls pydantic_ai.models.infer_model(self.model) and fails immediately if the provider implied by AGENT_MODEL (any provider:model string pydantic_ai recognizes — anthropic:, openai:, google:, ollama:, …) is misconfigured. This delegates to pydantic_ai's own provider classes rather than a hardcoded per-provider list, so a provider pydantic_ai adds in a future release is validated automatically with no changes needed here. Agent-specific env vars carry an AGENT_ prefix; provider API keys and LOGFIRE_TOKEN deliberately don't, because the provider SDKs read those standard names directly — Settings.model_config sets extra="ignore" specifically so an unprefixed, undeclared key like GOOGLE_API_KEY sitting in .env doesn't trip pydantic-settings' extra="forbid" default before validation even runs. This means import agent.config (or anything that imports it transitively) can fail before any code runs, which is the point — but it's also why every module under agent/ needs some valid provider config present at import time, even for code paths that never call the model.

  • Unit tests never need real credentials — three layers guarantee it. The root conftest.py forces AGENT_MODEL=test (Pydantic AI's built-in model) at import time unless the command line selects the live tests (-m eval), so the offline suite — and the first stage of scripts/release_check.py — passes whatever model and key .env configures, even with none (this has to happen at import, from sys.argv, because tests/conftest.py imports agent.config before any pytest hook runs; tests/test_hermetic.py guards it). tests/conftest.py calls os.environ.setdefault("ANTHROPIC_API_KEY", ...) / OPENAI_API_KEY before importing anything from agent/ (satisfying the import-time validator), and an autouse fixture overrides every Agent under agent.agents and examples — including nested worker agents — with TestModel (all agent modules are pre-imported so lazy imports can't escape it), so no unit test can ever hit a real model API. That TestModel is built with call_tools=[] on purpose: a default TestModel() calls every tool with junk arguments ("a"), which fails any tool that validates its input with ModelRetry and really executes tools that write, send or bill. So no tool runs in unit tests unless a test opts in with TestModel(call_tools=["name"]) (see smoke_tools in examples/supervisor/example.toml for the supervisor's delegation tool, and tests/test_safety_net.py for the recipe); test tool logic by calling the function directly. Don't remove either — together they're the reason uv run pytest works with zero setup and zero spend. evals/conftest.py deliberately does none of this: a missing key there should fail loudly, since evals make real API calls anyway.

  • There is no canonical agent; examples are the source of truth. agent/agents/ ships empty. Every pattern lives in examples/<name>/ (agent.py, prompts/, README.md, and an example.toml manifest read by scripts/example_manifest.py). tests/test_examples.py, tests/test_content_filter.py and tests/test_cost_limit.py are parametrized over every example, and tests/test_add_agent.py runs the real script into a scratch copy of the repo and tests what it generates — adding an example adds its checks. Examples run in place: examples/__init__.py registers each example's prompts/ directory with load_prompt (via PROMPTS_DIRS), and a copied agent finds its prompt in agent/prompts/. An example's prompt files are named <example>.txt or <example>_<role>.txt so add_agent.py can rename them (and rewrite the matching load_prompt("<example>...") calls) to the new agent's name. Only blank has its symbols renamed (blank is a placeholder token, templated = true); keep the word out of that file otherwise. Examples that need extra packages declare them in example.toml (dependencies), never in the root pyproject.toml; their tests skip when the packages are absent.

  • Every run_* helper returns a RunResult (agent/runs.py), never the bare output. It holds .output, the total .usage, and .steps — one Step(agent, result) per agent run, each holding the native AgentRunResult. A single agent is a one-step run; routers, pipelines and fan-outs record several. Build one with Flow: flow = Flow(USAGE_LIMITS), await flow.run(agent, prompt, deps=deps) for each step, return flow.finish(output). Flow owns one shared RunUsage, so USAGE_LIMITS bounds the whole flow, not each call (the steps share one usage object, which is why RunResult.usage is stored rather than summed). The generic tests drive each example's whole flow through run and assert on the result, so a new example is checked end to end for free. Don't make a helper return only the output: it hides usage and messages from callers, tests and the web UI.

  • Trace labels and smoke configuration. Every Agent is labeled with name=LABEL (LABEL = agent_label(__name__) from agent/logging.py; helpers use f"{LABEL}.role"), so Logfire traces carry the name you gave the agent, not the example's; tests/test_examples.py and tests/test_add_agent.py enforce it. Keep every Agent reachable from module scope (as a variable, or in a module-level dict/list): the TestModel safety net finds agents there and nowhere else. The offline smoke tests give each agent a TestModel that calls no tools and returns generated output. An agent that needs something else is configured in example.toml under [smoke.<agent_variable>]: call_tools = [...] opts in to calling tools (and then expects them called), and output = {...} is the output to return when its validators reject TestModel's generated junk (see examples/supervisor and examples/extraction). [entrypoint] names the deps class and the run helper.

  • Two models, deliberately different. settings.model (default anthropic:claude-sonnet-5-5) is the agent under test; settings.judge_model (default anthropic:claude-opus-5-5, in evals/judge.py) grades its output. They're kept separate to avoid self-assessment bias — but the judge should stay at least as capable as the agent, not cheaper/weaker, or the grading itself becomes the unreliable part.

  • Span-based evals depend on evals/conftest.py's setup_logging fixture. The agentic evaluators (MaxModelRequests, MaxToolCalls, ToolCorrectness, …) read the tool/request trajectory from OpenTelemetry spans, which only exist because that autouse fixture calls configure_logging() (Logfire + instrument_pydantic_ai()). Remove it and those evaluators fail with a "no span tree" reason rather than silently passing. Note pydantic_evals.Dataset requires a name=.

  • Tool error convention (three outcomes): ModelRetry (see agent/tools/example.py) is reserved for errors the LLM can plausibly fix by changing its input — bad query format, out-of-range params. ToolFailed is for expected, terminal failures the LLM can't fix but can work around — resource not found, unsupported operation: it returns the failure to the model without retry instructions and without spending the tool's retry budget (repeated failures are bounded by USAGE_LIMITS). Anything unexpected is logged and re-raised as a normal exception. Don't reach for ModelRetry as a generic catch-all (it burns the retry budget on failures the model can't correct), and don't wrap except Exception in ToolFailed (it hides bugs from the logs and the caller) — raise it only for failures you anticipated, with a message that says what to do instead.

  • Every run is bounded by USAGE_LIMITS. Exceeding request_limit, total_tokens_limit or (if set) cost_limit raises UsageLimitExceeded rather than silently looping. cost_limit is optional and off by default: every agent passes settings.cost_limit (AGENT_COST_LIMIT), which is None unless you set it. Leave it unset for models Pydantic AI can't price (e.g. ollama:) — their RunUsage.cost is None, so a set cap can't be enforced and Pydantic AI emits a CostNotFoundWarning on every process. With no cap set there is no warning (tests/test_cost_limit.py guards this), so don't re-add a test-time warning filter that would hide it. If an agent legitimately needs more iterations, raise the limit in your agent module — don't remove the guardrail.

  • Every agent carries RaiseContentFilterError. Without it pydantic_ai only raises when a content-filtered response is empty; a partial or refusal response would burn the output-retry budget re-prompting a refused request (structured output) or be returned as if complete (str). With it, any finish_reason='content_filter' response raises ContentFilterError out of run_* (the response is in .body). run_* deliberately doesn't catch it — like UsageLimitExceeded, callers decide what to show. In the supervisor a worker's ContentFilterError aborts the run rather than becoming a ModelRetry. Keep the capability on any agent you write (the examples and scripts/add_agent.py do).

  • Logfire falls back to console automatically when LOGFIRE_TOKEN is unset — there's no separate "dev mode" flag. If you're expecting cloud traces and only seeing console output, check .env for the token first. Traces are tagged from settings.service_name / settings.environment (AGENT_SERVICE_NAME, AGENT_ENVIRONMENT; an unset environment defers to LOGFIRE_ENVIRONMENT). AGENT_LOG_CONTENT=false strips prompts, outputs and tool arguments from spans, but evals/conftest.py calls configure_logging(include_content=True) regardless, because ArgumentCorrectness can't read tool arguments without them — don't remove that override.