Design notes¶
This repo is a template: a clean starting point for a production-quality Pydantic AI agent, meant to be cloned and reshaped into a specific agent. The code is the source of truth.
Commands¶
uv sync --group dev # install deps
uv run pytest # unit tests only (TestModel, no API key needed)
uv run pytest -m eval # all evals — pass/fail AND LLM-judge (needs a real API key, costs money)
uv run ruff check . # lint
uv run ruff format . # format
There is no separate llm_judge marker. Everything in evals/ carries only @pytest.mark.eval, so -m eval runs it all in one shot — there's no cheaper eval-only subset to reach for.
asyncio_mode = "auto" is set in pyproject.toml, so async tests need no @pytest.mark.asyncio decorator — don't add them back. The loop scopes are also pinned to session there: agents are module-level, so each provider HTTP client binds to the loop of the first test that uses it, and per-test loops (the default) make every later real model call — i.e. every eval after the first — fail with Event loop is closed. tests/test_event_loop.py guards this.
Changelog¶
CHANGELOG.md follows Keep a Changelog and is written for people who cloned the template. Any user-visible change (new or changed setting, default, convention, scaffold output, or anything a cloner must act on) adds a line under ## [Unreleased] in the same commit, grouped as Added / Changed / Fixed / Removed, with an "Upgrade notes" entry when action is required. To release: bump version in pyproject.toml (and uv lock), rename [Unreleased] to the new version with its date, add a fresh empty [Unreleased], update the compare links at the bottom, and tag vX.Y.Z.
Releasing¶
Before tagging, run uv run python scripts/release_check.py (Docker must be running if any example declares services) (makes real model calls: needs the provider key for AGENT_MODEL and costs money — pennies per example on a small model). It exists because TestModel can't tell you an example works: it ignores what tools return and what instructions say, so a placeholder tool or a half-written prompt passes every offline test (the tool_calling example once shipped a tool that echoed its query, and a real model looped on it until the request limit). The gate:
- runs the offline suite under coverage;
- runs each example in its own
uv runenvironment (only its declareddependencieslayered on), under a hard spend cap from itscost_budget_usd: first its live tests (examples/<name>/test_live.py, markedeval— real-model assertions on behavior, plus a requirement that every agent the example defines actually ran, read from spans viaevals/trace.pyso workers inside tool calls count, plus a run of the module'sif __name__ == "__main__"demo), then a smoke run (scripts/record_example.py: the run returns aRunResult, steps are labeled for the example, the manifest'sexpected_toolswere really called, cost stayed in budget); - fails unless every line of every
examples/*/agent.pywas executed by the offline and live tests together (coverage report,fail_under = 100inpyproject.toml) — so no example code is untested; - prints pass/fail/unverified counts, tokens and spend, and exits non-zero unless everything passed.
- when the whole library passed (all examples, tests included, nothing unverified), writes
badges/coverage.json, which the README's shields.io endpoint badge reads; it records the coverage percentage (the badge readscoverage: 100%) and is never written by a partial run (badge_is_earnedinscripts/release_check.py). Commit it with the release.
The gate checks whatever model AGENT_MODEL names, so you can reconfigure .env freely: the live tests assert behavior that holds for any capable model (tool outcomes, structure, which agents ran, no invented data), not one model's wording or request counts, and nothing in the gate names a provider. What a model must be able to do is the floor: call tools and return valid structured output. A model that can't (small local ones often return invalid JSON) fails the gate with Pydantic validation errors, which is a verdict on the model, not the example. The model that recorded each transcript is in its header, and --record overwrites them, so record with the model you want the docs to show.
--record also rewrites each example's sample_run.md (a transcript built from the RunResult; a failed run never overwrites the last good one — commit the result). scripts/record_example.py <name> records one example. Run a single example's live tests with uv run pytest -m eval examples/<name>. When you add an example, give it a test_live.py (copy the closest existing one): inputs with an unambiguous expected result, semantic assertions, every agent exercised, and the demo script. The gate is deliberately a manual, local step — it spends money and needs your key — so it is not wired into CI (.github/workflows/ci.yml runs only the offline suite). Run it yourself, with the provider key in .env or the environment; the model comes from AGENT_MODEL (the committed transcripts were recorded with google:gemini-3.1-flash-lite). An example that declares services is started with docker compose up --build --wait for the duration of its checks and always torn down; if Docker isn't usable it is reported as unverified with the reason (and fails the check unless --allow-unverified), never as passed.
Documentation site¶
The docs site (MkDocs Material, published to GitHub Pages by .github/workflows/docs.yml, served at https://tmtabor.io/agent-template/) is generated by docs/gen_pages.py from the repo's own files, with two hand-written exceptions. The sources: the README (split into pages by its ## sections; the home page is its header, pitch and sales sections), AGENTS.md (the "Design notes" page), CHANGELOG.md, CONTRIBUTING.md, MAINTAINING.md, each example (its README, sample_run.md, and its source in tabs), and the two guides in docs/pages/ ("Which pattern should I use?" and the FAQ), which are the only hand-written pages. Links are written repo-relative, like everything else, and rewritten to site pages or GitHub (including README.md#some-section, which lands on the page that now holds the section). So to change the docs, change those files. Consequences worth knowing: (1) the README's ## section titles are an interface: the ones in README_PAGES and HOME_SECTIONS in docs/gen_pages.py must exist, and renaming one makes the build fail loudly (SectionNotFound) until the generator is updated; (2) the README's "What are you building?" section is the home page's tabs: each ### under it becomes a tab on the site and an ordinary headed section on GitHub, so keep each one self-contained; (3) the README opens with a centred HTML header (logo in <picture>, tagline, badges, links): the site swaps the <picture> for Material's #only-light / #only-dark images, drops the links back to the site itself, and turns raw-HTML .md links into the directory URLs MkDocs serves, because MkDocs only fixes Markdown links; (4) the logo SVGs and the favicon are in docs/assets/ (the wordmark is drawn as outlines, so it needs no font); (5) a new example gets its page and nav entry automatically, needs a sample_run.md (tests/test_examples.py enforces it) and optionally a slot in DISPLAY_ORDER in scripts/example_manifest.py; (6) examples/README.md is generated (uv run python scripts/examples_index.py) and tests/test_docs.py fails if it is stale; (7) docs/pages/ is not published as it is, and on a pattern page the recorded run is a closed collapsible block below the source, titled with the run's model, steps and cost, behind an anchor that docs/assets/extra.js opens when a link points at #recorded-run (the README's "See it run" line); (8) write for GitHub but expect MkDocs: the generator adds the blank line MkDocs needs before a list that follows a **bold** line, but anything else GitHub forgives and MkDocs doesn't should be checked by looking at the page (mkdocs serve), because --strict only catches broken links and a raw-HTML link is not checked at all. The site builds with mkdocs build --strict, so a broken link fails the build, and tests/test_docs.py checks every internal link and nav entry without needing MkDocs. mkdocs is pinned <2: MkDocs 2.0 will break plugins and themes, and properdocs (a continuation of 1.x that Material already pulls in) is the drop-in if you need to leave 1.x. The docs dependency group is separate from dev; add_agent.py --prune removes the docs, their tests, that group, and CONTRIBUTING.md and MAINTAINING.md. The how-to for maintainers is in MAINTAINING.md.
Writing a new example¶
Everything a pattern example must have, because the gate (see Releasing) checks all of it and a TestModel can hide a broken example from every offline test:
agent.pywith arun_*helper returning aRunResult, a deps dataclass constructible with no arguments, andname=LABELon everyAgent(f"{LABEL}.role"for helpers) — keep every agent reachable from module scope. No placeholder tools, instructions or data: a real model will loop on a tool that echoes its input. Make the domain invented or deterministic, so a correct answer proves the pattern worked (RAG's fictional store policies, MCP's date arithmetic).example.toml:smoke_inputthat makes a real model exercise the pattern,expected_tools, and a[smoke.<agent>]table for any agent whose validators rejectTestModel's junk (output = {...}), whose bool defaults the flow needs flipped, or that must call tools. Adependenciesentry for anything the rootpyproject.tomllacks.test_example.py: offline tests with scriptedFunctionModels covering every branch — the gate requires 100% line coverage ofagent.pyover offline and live tests together. Test the side effect, not the model's words (the refund ledger, the retrieved ids), and any ground truth the example defines.test_live.py(markedeval): inputs with an unambiguous expected result, assertions that hold for any capable model (no exact request counts, no one model's phrasing), a requirement that every agent ran (assert_every_agent_ran), and a run of the__main__demo. When the example has a pause, a block or a rejection, observe it in the spans or the ground truth, not just the final text.README.md(with the**See it run:**line), added toDISPLAY_ORDERinscripts/example_manifest.py, thenuv run python scripts/examples_index.py, thenuv run python scripts/release_check.py <name> --record.
An example with dependencies or test_dependencies is skipped by the generic tests in the default environment (an ImportError from a declared dependency is a skip, not a failure) and runs, with the generic and add_agent tests, in its own uv run --with environment inside the gate. dependencies are what using the agent needs (add_agent.py installs them); test_dependencies are what only its tests need (for instance the server half of a library, to run a service locally). If a library raises a plain ImportError when an extra is missing (as fastmcp does for its server half), probe for that half in the test file rather than the package name.
An example that runs untrusted code (code_mode is the model) must test the sandbox with real hostile input, not only the happy path, and check the host, not the sandbox's own error message: that no file appeared, no secret came back, no tool ran beyond its cap, and a runaway loop stopped at its time limit. Assert on the invariant rather than Monty's wording (it has no open at all until os is imported, then raises PermissionError). Monty is a Python subset, and it type-checks the model's code against the tool signatures before running it: for example next() over a generator expression fails, and x = None later passed as a str is rejected. Put what the model needs to avoid those in the prompt or the tool types (Literal categories stopped it comparing against 'Travel'), and measure the effect in model requests. pydantic-ai-harness is pre-1.0 and its minor version tracks Pydantic AI's (0.54 goes with 2.54), so bump them together; the dependencies pin is >=0.54,<1.
A planner-executor example (planner_executor is the model) must enforce the plan's rules in code and test each rule with a plan that breaks it, because a real model does not reliably follow them: Gemini wrote "compare them" steps with depends_on left empty in 9 of 10 plans, even with the rule spelled out and an example in the prompt, and the answers still came out right (the executors had the lookup tool and the synthesizer did the arithmetic), so a correct answer proves nothing about the plan. Check the plan's shape in the live tests (independent steps in one wave, a step that depends on others), not only the answer. The text a rejected plan is sent back with is part of the prompt: saying what was wrong made the model repeat the same plan three times; adding what to change (computed by the code) fixed it on the first retry in 11 of 12 runs. Prove the scheduling with scripted models that record what each executor was told and how many ran at once, and mutate the implementation (run steps one at a time, leak every result to every executor, stop the skip cascade early) to see that the tests notice.
A durable example (temporal is the model) must be tested against the real Temporal server, because Temporal's own test server downloads a binary at run time and a hermetic suite cannot: its tests skip without the service's address and the release check supplies it. Assert on the server's event history (fetch_history_events), not on the model's words: the model-request activity must show one start per request, and a retried tool shows only its final attempt number (failed attempts of a retried activity are not events), so the failures themselves belong in a ledger your own tool keeps. Prove crash recovery with a real SIGKILL of a worker in a separate process as well as a tidy shutdown (the server notices a dead worker only when the activity's timeout passes). Things that bite: a workflow cannot be defined in __main__, so the demo block re-imports the module under its real name; the agent's name (and therefore its activity names) must be the same in the worker and in the workflow's sandbox; the template's own modules must be imported inside workflow.unsafe.imports_passed_through() because the sandbox forbids what they do at import time; two workflows cannot share an agent; a bug in workflow code makes Temporal retry forever, so every run needs an execution timeout; and Pydantic AI's invoke_agent span does not appear for a run inside a workflow, so use the history as the evidence that the agent ran.
An example that needs a running service (mcp_tools is the model; temporal is the image-only variant) ships service/ (the server and a Dockerfile, or just a compose file naming a pinned published image, never latest; a docker-compose.yml that publishes the port as "127.0.0.1::<port>" and has a healthcheck, so up --wait returns when it is ready and nothing is exposed beyond this machine) and declares it in example.toml (services, and a [service.<name>] table with the container port, the env variable that receives its address, and a url template). The agent reads the address from deps, so tests and the gate can point it anywhere. Three layers of testing, none needing Docker until the last: the server's functions directly; the server run as a local subprocess on a free port for the offline tests (a fixture in the example's conftest.py); and the real container, which only the release check starts. Generic tests that run the agent skip such an example unless its address variable is set (import_example(example, running=True)), so the default suite never needs Docker.
Making it yours¶
The intended customization sequence, roughly in order:
- Add an agent:
uv run python scripts/add_agent.py(menu) oradd_agent.py <example> --name <name>. The template ships with no agents. The script copies an example fromexamples/intoagent/agents/<name>.py, its prompt intoagent/prompts/, and scaffoldstests/test_agents_<name>.pyandevals/test_<name>.py. Run it once per agent; each can use a different pattern. Import agents directly from their own modules — there is no shared "primary" agent. - Edit the prompt:
agent/prompts/<name>.txt, loaded viaload_prompt("<name>"). Add more.txtfiles beside it and load them the same way. - Define the output schema: replace the placeholder fields on the output model in your agent module. The examples keep a
result: strfield, which the eval starters and the web-UI skill read when present (evals/helpers.py:output_textfalls back tostr(output)). Keep the schema as flat as the data actually requires — don't wrap a single field in its own object (e.g. preferlist[str]overlist[{text: str}]) unless a second field genuinely needs to travel with it. This matters more the smaller/weakersettings.modelis: a schema a frontier model satisfies without issue can reliably burn the retry budget on a small local model (e.g. anollama:model) if it adds nesting the data doesn't need. If output validation keeps failing against the configured model, check whether the schema is more nested than necessary before assuming it's a prompting problem. - Add tools: copy the pattern in
agent/tools/example.py, register with@<your_agent>.tool. - Tune
USAGE_LIMITSin your agent module:request_limitcaps model round-trips per run,total_tokens_limitcaps tokens. An optional USD spend cap comes fromAGENT_COST_LIMIT(off by default). The supervisor shares its budget with workers viausage=ctx.usage. - Grow the evals: add cases to
evals/fixtures/<name>.json(picked up by the dataset eval inevals/test_<name>.pyautomatically) and adapt the judge criteria there. Fixtures also accept optionalexpected_tools,expected_trajectoryandexpected_argumentskeys, which become span-based behavioral evaluators (ToolCorrectness,TrajectoryMatch,ArgumentCorrectness); every case getsMaxModelRequests/MaxToolCallsbudgets. The shared evaluators and runner are inevals/helpers.py. - Drop what you don't need:
uv run python scripts/add_agent.py --prunedeletes the examples you didn't use (keepingblank), the docs and their tests. Nothing underagent/orevals/depends onexamples/.
Non-obvious architecture¶
-
agent/config.pyvalidates at import time, not at call time.Settingshas amodel_validatorthat callspydantic_ai.models.infer_model(self.model)and fails immediately if the provider implied byAGENT_MODEL(anyprovider:modelstring pydantic_ai recognizes —anthropic:,openai:,google:,ollama:, …) is misconfigured. This delegates to pydantic_ai's own provider classes rather than a hardcoded per-provider list, so a provider pydantic_ai adds in a future release is validated automatically with no changes needed here. Agent-specific env vars carry anAGENT_prefix; provider API keys andLOGFIRE_TOKENdeliberately don't, because the provider SDKs read those standard names directly —Settings.model_configsetsextra="ignore"specifically so an unprefixed, undeclared key likeGOOGLE_API_KEYsitting in.envdoesn't trip pydantic-settings'extra="forbid"default before validation even runs. This meansimport agent.config(or anything that imports it transitively) can fail before any code runs, which is the point — but it's also why every module underagent/needs some valid provider config present at import time, even for code paths that never call the model. -
Unit tests never need real credentials — three layers guarantee it. The root
conftest.pyforcesAGENT_MODEL=test(Pydantic AI's built-in model) at import time unless the command line selects the live tests (-m eval), so the offline suite — and the first stage ofscripts/release_check.py— passes whatever model and key.envconfigures, even with none (this has to happen at import, fromsys.argv, becausetests/conftest.pyimportsagent.configbefore any pytest hook runs;tests/test_hermetic.pyguards it).tests/conftest.pycallsos.environ.setdefault("ANTHROPIC_API_KEY", ...)/OPENAI_API_KEYbefore importing anything fromagent/(satisfying the import-time validator), and an autouse fixture overrides everyAgentunderagent.agentsandexamples— including nested worker agents — withTestModel(all agent modules are pre-imported so lazy imports can't escape it), so no unit test can ever hit a real model API. ThatTestModelis built withcall_tools=[]on purpose: a defaultTestModel()calls every tool with junk arguments ("a"), which fails any tool that validates its input withModelRetryand really executes tools that write, send or bill. So no tool runs in unit tests unless a test opts in withTestModel(call_tools=["name"])(seesmoke_toolsinexamples/supervisor/example.tomlfor the supervisor's delegation tool, andtests/test_safety_net.pyfor the recipe); test tool logic by calling the function directly. Don't remove either — together they're the reasonuv run pytestworks with zero setup and zero spend.evals/conftest.pydeliberately does none of this: a missing key there should fail loudly, since evals make real API calls anyway. -
There is no canonical agent; examples are the source of truth.
agent/agents/ships empty. Every pattern lives inexamples/<name>/(agent.py,prompts/,README.md, and anexample.tomlmanifest read byscripts/example_manifest.py).tests/test_examples.py,tests/test_content_filter.pyandtests/test_cost_limit.pyare parametrized over every example, andtests/test_add_agent.pyruns the real script into a scratch copy of the repo and tests what it generates — adding an example adds its checks. Examples run in place:examples/__init__.pyregisters each example'sprompts/directory withload_prompt(viaPROMPTS_DIRS), and a copied agent finds its prompt inagent/prompts/. An example's prompt files are named<example>.txtor<example>_<role>.txtsoadd_agent.pycan rename them (and rewrite the matchingload_prompt("<example>...")calls) to the new agent's name. Onlyblankhas its symbols renamed (blankis a placeholder token,templated = true); keep the word out of that file otherwise. Examples that need extra packages declare them inexample.toml(dependencies), never in the rootpyproject.toml; their tests skip when the packages are absent. -
Every
run_*helper returns aRunResult(agent/runs.py), never the bare output. It holds.output, the total.usage, and.steps— oneStep(agent, result)per agent run, each holding the nativeAgentRunResult. A single agent is a one-step run; routers, pipelines and fan-outs record several. Build one withFlow:flow = Flow(USAGE_LIMITS),await flow.run(agent, prompt, deps=deps)for each step,return flow.finish(output).Flowowns one sharedRunUsage, soUSAGE_LIMITSbounds the whole flow, not each call (the steps share one usage object, which is whyRunResult.usageis stored rather than summed). The generic tests drive each example's whole flow throughrunand assert on the result, so a new example is checked end to end for free. Don't make a helper return only the output: it hides usage and messages from callers, tests and the web UI. -
Trace labels and smoke configuration. Every
Agentis labeled withname=LABEL(LABEL = agent_label(__name__)fromagent/logging.py; helpers usef"{LABEL}.role"), so Logfire traces carry the name you gave the agent, not the example's;tests/test_examples.pyandtests/test_add_agent.pyenforce it. Keep every Agent reachable from module scope (as a variable, or in a module-level dict/list): theTestModelsafety net finds agents there and nowhere else. The offline smoke tests give each agent aTestModelthat calls no tools and returns generated output. An agent that needs something else is configured inexample.tomlunder[smoke.<agent_variable>]:call_tools = [...]opts in to calling tools (and then expects them called), andoutput = {...}is the output to return when its validators rejectTestModel's generated junk (seeexamples/supervisorandexamples/extraction).[entrypoint]names the deps class and therunhelper. -
Two models, deliberately different.
settings.model(defaultanthropic:claude-sonnet-5-5) is the agent under test;settings.judge_model(defaultanthropic:claude-opus-5-5, inevals/judge.py) grades its output. They're kept separate to avoid self-assessment bias — but the judge should stay at least as capable as the agent, not cheaper/weaker, or the grading itself becomes the unreliable part. -
Span-based evals depend on
evals/conftest.py'ssetup_loggingfixture. The agentic evaluators (MaxModelRequests,MaxToolCalls,ToolCorrectness, …) read the tool/request trajectory from OpenTelemetry spans, which only exist because that autouse fixture callsconfigure_logging()(Logfire +instrument_pydantic_ai()). Remove it and those evaluators fail with a "no span tree" reason rather than silently passing. Notepydantic_evals.Datasetrequires aname=. -
Tool error convention (three outcomes):
ModelRetry(seeagent/tools/example.py) is reserved for errors the LLM can plausibly fix by changing its input — bad query format, out-of-range params.ToolFailedis for expected, terminal failures the LLM can't fix but can work around — resource not found, unsupported operation: it returns the failure to the model without retry instructions and without spending the tool's retry budget (repeated failures are bounded byUSAGE_LIMITS). Anything unexpected is logged and re-raised as a normal exception. Don't reach forModelRetryas a generic catch-all (it burns the retry budget on failures the model can't correct), and don't wrapexcept ExceptioninToolFailed(it hides bugs from the logs and the caller) — raise it only for failures you anticipated, with a message that says what to do instead. -
Every run is bounded by
USAGE_LIMITS. Exceedingrequest_limit,total_tokens_limitor (if set)cost_limitraisesUsageLimitExceededrather than silently looping.cost_limitis optional and off by default: every agent passessettings.cost_limit(AGENT_COST_LIMIT), which isNoneunless you set it. Leave it unset for models Pydantic AI can't price (e.g.ollama:) — theirRunUsage.costisNone, so a set cap can't be enforced and Pydantic AI emits aCostNotFoundWarningon every process. With no cap set there is no warning (tests/test_cost_limit.pyguards this), so don't re-add a test-time warning filter that would hide it. If an agent legitimately needs more iterations, raise the limit in your agent module — don't remove the guardrail. -
Every agent carries
RaiseContentFilterError. Without it pydantic_ai only raises when a content-filtered response is empty; a partial or refusal response would burn the output-retry budget re-prompting a refused request (structured output) or be returned as if complete (str). With it, anyfinish_reason='content_filter'response raisesContentFilterErrorout ofrun_*(the response is in.body).run_*deliberately doesn't catch it — likeUsageLimitExceeded, callers decide what to show. In the supervisor a worker'sContentFilterErroraborts the run rather than becoming aModelRetry. Keep the capability on any agent you write (the examples andscripts/add_agent.pydo). -
Logfire falls back to console automatically when
LOGFIRE_TOKENis unset — there's no separate "dev mode" flag. If you're expecting cloud traces and only seeing console output, check.envfor the token first. Traces are tagged fromsettings.service_name/settings.environment(AGENT_SERVICE_NAME,AGENT_ENVIRONMENT; an unset environment defers toLOGFIRE_ENVIRONMENT).AGENT_LOG_CONTENT=falsestrips prompts, outputs and tool arguments from spans, butevals/conftest.pycallsconfigure_logging(include_content=True)regardless, becauseArgumentCorrectnesscan't read tool arguments without them — don't remove that override.