Skip to content

Agents

The template starts with no agents. When you want one, run add_agent.py and pick a pattern from the menu; run it again for each further agent, each with whatever pattern suits it.

uv run python scripts/add_agent.py                       # interactive menu
uv run python scripts/add_agent.py supervisor --name triage
uv run python scripts/add_agent.py blank --name newsletter

The patterns live in examples/ — each is a folder with the agent's source, its prompt, a README and an example.toml:

Example Pattern
blank An empty agent: one output type, one prompt, no tools
single One agent handles the whole task
supervisor A supervisor delegates to specialized workers
planner_executor A planner writes the whole plan as data; code checks it and runs it, independent steps in parallel; a last agent answers
tool_calling An agent whose tools call external systems
extraction Free text to a validated schema, with an output validator and retry budget
rag Answer from your own documents by meaning, with embeddings in a Chroma vector database (a Docker service; needs chromadb-client), and cite only what was really retrieved
mcp_tools Use the tools of an MCP server that runs as its own Docker service
code_mode The model writes Python that calls your tools in a sandbox (Monty): exact answers from one or two requests (needs pydantic-ai-harness)
temporal A durable agent run as a Temporal workflow: failing tools are retried and a crashed worker is replaced, without repeating model calls (Temporal runs as a Docker service; needs temporalio)
conversation Memory across turns, a bounded context window, and streaming
human_in_the_loop Pause a risky tool call for approval, reject impossible ones first, resume the run
guardrails Check input in code and with a guard model, validate output, turn failures into safe answers
router A classifier picks a category; code dispatches to a specialist
pipeline Fixed sequential steps, each output feeding the next, with gates
fan_out Parallel workers via asyncio.gather, then an aggregator
evaluator_optimizer A generator and a critic loop until the output passes or a cap is hit

For each agent, add_agent.py:

  • copies the example's module to agent/agents/<name>.py and its prompt to agent/prompts/<name>.txt (only blank renames its symbols; the others keep theirs, which is safe because each agent lives in its own module),
  • scaffolds a smoke test, tests/test_agents_<name>.py (runs under TestModel, no API key),
  • scaffolds an eval starter, evals/test_<name>.py, with a fixture file at evals/fixtures/<name>.json,
  • copies the example's service, if it has one, to services/<name>/ (mcp_tools ships its MCP server, a Dockerfile and a compose file; temporal a compose file for the Temporal server) and tells you how to start it,
  • runs uv add for any extra dependencies the example declares (code_mode needs pydantic-ai-harness[code-mode], temporal needs temporalio, rag needs chromadb-client), and tells you about any environment variables it needs.

There is no shared "primary" agent. Import each agent directly from its own module. Every run_* helper returns the same thing, a RunResult (agent/runs.py):

from agent.agents.triage import run_supervisor

result = await run_supervisor("Summarize the benefits of unit tests")
result.output  # the validated output
result.usage  # total usage, across every agent run in the flow
result.steps  # each agent run, in order: Step(agent="triage", result=<AgentRunResult>)

A single agent is a one-step run, a router or pipeline a several-step run, so moving an agent from one shape to the other never changes a call site. result.steps[0].result is the native Pydantic AI result if you want its messages or run id.

Once you've picked what you need, uv run python scripts/add_agent.py --prune removes the other examples (keeping blank), the docs and their tests. Nothing under agent/ or evals/ depends on examples/.