Changelog¶
All notable changes to this template are documented here. The format follows Keep a Changelog, and versions follow Semantic Versioning. Entries are written for people who have cloned the template: what changed, and whether you need to do anything.
Unreleased¶
0.3.0 - 2026-10-07¶
Version 0.3.0 turns the template from "one agent plus three stubs" into a library of seventeen working patterns that you add to your own project with one command, each run against a real model, with a documentation site.
Upgrade notes¶
- The template no longer ships an agent.
agent/agents/is empty and there is no canonicalrun_agent/AgentOutput/AgentDeps/agentre-export. Runuv run python scripts/add_agent.py(once per agent) and import each agent from its own module, e.g.from agent.agents.triage import .... Code importing the canonical names fromagent.agentsmust change. scripts/choose_pattern.pyis gone.add_agent.pyreplaces it and the oldadd_agent.py <name>scaffold:add_agent.py supervisor --name triage, oradd_agent.py blank --name newsletterfor what the old script produced. Existing clones are unaffected until you pull this change; if you already chose a pattern, your agent keeps working butagent/agents/__init__.pyno longer needs the canonical import.evals/test_pass_fail.py,evals/test_llm_judge.pyandevals/fixtures/example.jsonare removed. Each agent now getsevals/test_<name>.pyandevals/fixtures/<name>.jsonfromadd_agent.py; the shared evaluators moved toevals/helpers.py. Copy your fixtures to the new file name.agent/prompts/system.txtis gone. Each agent hasagent/prompts/<name>.txt.run_*helpers return aRunResult, not the bare output. Read.outputfor what you used to get back:result = await run_agent(...), thenresult.output.result.usageis the total usage andresult.stepsrecords each agent run (result.steps[0].resultis the native Pydantic AI result). Applies torun_agent,run_tool_agent,run_supervisorand every new example's helper; agents you copy withadd_agent.pyget the same shape.- The
agent-web-uiskill'schat.pyimports one agent module you point it at (see the skill's "Before you start"), instead of the canonical names. - Python 3.13 is still the minimum, and 3.14 is now supported. Python 3.15 is not yet.
- Some examples read their own environment variables (an embedding model, the address of a service). They are
optional, and only matter if you add that example; see "Configuration" in the README and
.env.example.
Added¶
- Seventeen example patterns in
examples/, each with its source, prompt, README, offline and live tests, and a recorded run against a real model: - Basics:
blank,single,conversation(message history, a bounded window, streaming). - Tools and data:
tool_calling,extraction(an output validator and a retry budget),rag(documents embedded and searched by meaning in Chroma, with citations checked against what was retrieved),mcp_tools(the tools of an MCP server running as a Docker service) andcode_mode(the model writes Python that calls your tools, in a sandbox). - Several agents:
supervisor,planner_executor(a plan written as data, checked and run by code),router,pipeline,fan_outandevaluator_optimizer. - Safety and reliability:
human_in_the_loop(pause a risky action for approval),guardrails(checks on input and output, failures turned into safe answers) andtemporal(a durable workflow: failing tools are retried, a dead worker is replaced). scripts/add_agent.py, the one way to add an agent: an interactive menu, oradd_agent.py <example> --name <name>. It copies the example and its prompt, scaffolds a smoke test and an eval starter, installs the example's extra packages, copies its service (if it has one) toservices/<name>/, and tells you which environment variables to set.--pruneremoves the examples you did not use, and the docs and release tooling with them.RunResultandFlow(agent/runs.py): everyrun_*helper returns the output, the total usage and one step per agent run, whatever the pattern, so a call site does not change when an agent grows from one step to several.- A documentation site at https://tmtabor.io/agent-template/, generated from the README, the examples and the other documents, so nothing is written twice. Each pattern has a page with an explanation, when to use it and when not to, its source, and its recorded run. Two guides: "Which pattern should I use?" and an FAQ.
CONTRIBUTING.md,MAINTAINING.mdandSECURITY.md(private vulnerability reporting is on), a logo, and a coverage badge in the README.--pruneremoves them, since they are about the template and not your project.- Python 3.14 support:
pyproject.tomllists it, CI runs lint and the offline suite on 3.13 and 3.14, andtests/test_python_versions.pykeeps the README badge, the classifiers and the CI matrix in agreement. - A release check (
scripts/release_check.py) for maintainers: the offline suite, then each example against a real model in its own environment (its live tests, a smoke run, its Docker service started and stopped for it), a gate that every line of every example's code ran, and the recorded runs. It is manual and local, and is not part of CI. - Examples can need more than the template.
example.tomldeclares extradependencies(installed only when you add that example),test_dependencies,services(adocker-compose.ymlthe release check starts), smoke-test overrides,expected_toolsand acost_budget_usd. - Smaller additions:
agent_label(__name__)labels each agent's traces with its own name (triage,triage.worker);evals/trace.pyreports every agent and tool call a run made, from its spans;evals/helpers.pyholds the shared evaluators and runner;Flow.runcan resume a run with no new prompt (message_history=,deferred_tool_results=);load_promptsearchesPROMPTS_DIRSso examples run in place.
Changed¶
- A new README and documentation home. They open with a logo, the line "Pick a pattern, edit the prompt, ship it.",
badges and a pitch, then a three-step "Get started", the use cases, the four opinions baked in, and next steps. The
maintainer sections moved to
MAINTAINING.md. tool_callingandsupervisornow do real work. Their placeholder tool and worker were replaced (a real model looped on the old echoing tool until it hit the request limit):tool_callinglooks up Python release notes with a tool that shows all three error outcomes, andsupervisorcoordinates a real analyst and writer.- The three pattern stubs moved from
agent/agents/toexamples/and are copied into your project on demand. - The unit-test safety net also covers
examples/and agents held in module-level dicts, lists and tuples, andtests/test_safety_net.pyno longer depends on a chosen agent. - The GitHub Actions in
.github/workflows/are on their current major versions (Node 24).
Fixed¶
- The offline test suite is hermetic. A root
conftest.pyforcesAGENT_MODEL=testunless the command line selects the live tests (-m eval), souv run pytestpasses with no provider key whatever.envconfigures. Before, it crashed on import if.envnamed a provider whose key was missing.
Removed¶
scripts/choose_pattern.py, the canonical re-export inagent/agents/__init__.py, andtests/test_stubs.py(replaced bytests/test_examples.py).
0.2.0 - 2026-10-04¶
Upgrade notes¶
CLAUDE.mdis nowAGENTS.md. If you customized it, move your edits over. Tools that look forCLAUDE.mdneed an@AGENTS.mdimport or a copy.- Evals:
pydantic_evals.Dataset(...)now requiresname=. Add it to any dataset you wrote;evals/test_pass_fail.pyalready does. - Unit tests no longer call agent tools by default. If you wrote tests that relied on
tools being called automatically, opt in with
TestModel(call_tools=["name"])(recipe intests/test_safety_net.py). - If you copied the pytest config, add
asyncio_default_test_loop_scope = "session"andasyncio_default_fixture_loop_scope = "session"; without them evals fail after the first. - Callers of
run_*should handleContentFilterError(see Added). It propagates instead of being retried or returned half-finished. - Default models changed to Sonnet 5.5 (agent) and Opus 5.5 (judge). Pin
AGENT_MODEL/AGENT_JUDGE_MODELif you want the old ones.
Added¶
- Optional
AGENT_COST_LIMIT(USD, e.g.0.50), off by default. When set it is passed ascost_limitin every stub's and scaffolded agent'sUSAGE_LIMITS. It is only enforceable for models with known pricing; leave it unset for others (e.g.ollama:), whose cost is unknown. With no limit set, Pydantic AI emits no cost warning. ToolFailedin the tool error convention: raise it for expected, terminal failures the model can work around (not found, unsupported). UnlikeModelRetryit spends no retry budget.agent/tools/example.pydemonstrates it.RaiseContentFilterErroron every agent: a content-filtered response now raisesContentFilterErrorout ofrun_*. The web-UI skill's chat router handles it.- Span-based agentic evals: every dataset case gets
MaxModelRequests/MaxToolCallsbudgets, and fixtures accept optionalexpected_tools,expected_trajectoryandexpected_argumentskeys (ToolCorrectness,TrajectoryMatch,ArgumentCorrectness). - Logfire settings:
AGENT_SERVICE_NAME(defaultagent; rename per project),AGENT_ENVIRONMENT(falls back toLOGFIRE_ENVIRONMENT) andAGENT_LOG_CONTENT(defaulttrue;falsekeeps prompts, outputs and tool arguments out of traces). Evals force content on becauseArgumentCorrectnessneeds tool arguments. CHANGELOG.md, and a convention inAGENTS.mdfor keeping it current.
Changed¶
- Upgraded to Pydantic AI 2.54.0 (from 2.3.0),
pydantic-evals2.54.0, Logfire 5.1.1,pydantic-settings2.15.0,python-dotenv1.2.4, pytest 9.1.1, pytest-asyncio 1.4.0 and ruff 0.16.10. Minimum versions inpyproject.tomlwere raised to match. - Default models:
anthropic:claude-sonnet-5-5for the agent andanthropic:claude-opus-5-5for the judge. agent/tools/example.pysearch now goes through a replaceable_search_backend().- ruff 0.16 formats Python code blocks in Markdown; the README was reformatted accordingly.
- Unit tests no longer call agent tools by default. The
tests/conftest.pysafety net now usesTestModel(call_tools=[]); a defaultTestModel()called every tool with junk arguments, which failed any tool that validates input withModelRetryand really ran tools with side effects. Tests that want a tool executed opt in withTestModel(call_tools=["name"])(seeSMOKE_TOOLSintests/test_stubs.pyand the recipe intests/test_safety_net.py). If you wrote tests that relied on tools being called automatically, add the opt-in.
Fixed¶
test_fixture_datasetcrashed underpydantic-evals2.54 becauseDatasetrequires aname(see Upgrade notes).- Evals made with a real model failed with
RuntimeError: Event loop is closedon every test after the first. Agents are module-level, so the provider's HTTP client was bound to the first test's event loop, which pytest-asyncio then closed.pyproject.tomlnow pinsasyncio_default_test_loop_scopeandasyncio_default_fixture_loop_scopetosession; if you copied the old pytest config, add both settings. Found while running the evals against a local Ollama model.
0.1.0 - 2026-08-06¶
Initial release: an opinionated starting point for a production-quality Pydantic AI agent.
Added¶
- Three agent patterns (
single,supervisor,tool_calling) behind a canonical import inagent/agents/__init__.py, switched withscripts/choose_pattern.py, plusscripts/add_agent.pyfor additional independent agents. - Import-time
Settingsvalidation that works for any pydantic-ai provider, with.envloading so keys set only in.envreach the provider SDKs. USAGE_LIMITSguardrail (request and token limits) on every run.- Logfire observability with automatic console fallback.
- Unit tests that run on
TestModelwith no credentials; evals (pass/fail dataset and LLM-as-judge) with separate agent and judge models. - Skills for scaffolding a web UI (
agent-web-ui) and a double-clickable macOS launcher (macos-launcher), plus CI (ruff and unit tests) and a BSD-3-Clause license.