Skip to content

Supervisor / workers

A supervisor agent decides which specialized workers to call, in what order, and what to hand each, then synthesizes the results.

A supervisor is an agent whose tools are other agents. Each worker has its own instructions, tools and output type. The supervisor sees them as delegation tools and decides, one turn at a time, which to call, in what order, and what to give each. It can call one worker, several, or none, and it writes the final answer from what they return. Nothing about the sequence is decided in advance. That is what makes it flexible, and also what makes it the least predictable of the multi-agent patterns.

Use it when

  • A task splits into specialised subtasks.
  • Workers need different tools, instructions or output types.
  • You want a coordinator that makes the routing decision itself, and you can't know the steps in advance.

Look elsewhere when

  • You can list the steps today: pipeline.
  • The input decides which specialist runs: router.
  • You want the plan checked before anything runs: planner_executor.
supervisor_agent → decides which worker to call, and what to hand it
  analyst_agent  → researches a question and returns key findings
  writer_agent   → turns material into clear prose

What it shows

  • Delegation as tools on the supervisor (delegate_to_analyst, delegate_to_writer): the model chooses whether to call one worker, both, or neither, and passes the analyst's findings to the writer itself
  • usage=ctx.usage so workers spend from the supervisor's shared budget
  • One USAGE_LIMITS bounding the whole delegation tree
  • The result has one step (the supervisor's); the workers ran inside it, so their calls are in .all_messages() and their spend is in .usage

See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.

uv run python scripts/add_agent.py supervisor --name my_agent

To adapt it, replace the two workers with your own and describe in the supervisor's instructions when each is the right call.

Source

All of it is in examples/supervisor/.

"""Supervisor/worker multi-agent pattern.

Use this pattern when:
- A task can be broken into specialized subtasks
- Different agents have different tools, instructions, or output types
- You want a coordinator that *decides* which workers to call, and in what order

Architecture:
    supervisor_agent → decides which worker to call, and what to hand it
    analyst_agent    → researches a question and returns key findings
    writer_agent     → turns material into clear prose

The supervisor sees each worker as a tool. The model chooses whether to call the analyst, the
writer, both (and in which order), or neither — that choice is what separates this from
`pipeline` (fixed order, chosen by code) and `router` (one specialist, chosen by code).

To adapt it:
    1. Replace the workers with your own, each with its specialized tools and instructions
    2. Give the supervisor one delegation tool per worker
    3. Describe, in the supervisor's instructions, when each worker is the right call
"""

from __future__ import annotations

from dataclasses import dataclass

from pydantic import BaseModel
from pydantic_ai import Agent, RunContext
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits

from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.runs import Flow, RunResult

logger = get_logger(__name__)
LABEL = agent_label(__name__)  # names this agent's run spans in Logfire traces

# Guardrail against runaway agentic loops. A run that exceeds any limit
# raises UsageLimitExceeded instead of silently burning tokens. Worker runs
# share the supervisor's budget (see the delegation tools), so this bounds
# the whole delegation tree, not just the supervisor's own requests. Set
# AGENT_COST_LIMIT (USD) to add a spend cap — optional, off by default, and only
# useful for models with known pricing (see Settings.cost_limit).
USAGE_LIMITS = UsageLimits(
    request_limit=12, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)


# --- Shared dependencies ---
@dataclass
class SharedDeps:
    """Dependencies shared across supervisor and worker agents."""

    pass


# --- Worker agents ---
# Each worker is a specialized agent with its own instructions and tools.


class Findings(BaseModel):
    points: list[str]


analyst_agent: Agent[SharedDeps, Findings] = Agent(
    settings.model,
    name=f"{LABEL}.analyst",  # helpers are labeled <agent>.<role>
    output_type=Findings,
    deps_type=SharedDeps,
    # Fail fast when the provider filters a response, instead of retrying a
    # refused request or returning partial text.
    capabilities=[RaiseContentFilterError()],
    instructions=(
        "You are an analyst. Given a question or topic, list its key findings as three to five "
        "short, factual points. Do not write prose; the points are handed to a writer."
    ),
)


class Draft(BaseModel):
    text: str


writer_agent: Agent[SharedDeps, Draft] = Agent(
    settings.model,
    name=f"{LABEL}.writer",
    output_type=Draft,
    deps_type=SharedDeps,
    capabilities=[RaiseContentFilterError()],
    instructions=(
        "You are a writer. Turn the material you are given into clear, concise prose that "
        "follows any instructions about length or audience. Use only the material provided."
    ),
)


# --- Supervisor agent ---
class SupervisorOutput(BaseModel):
    # `result` is the conventional output field in these examples; the generated
    # eval starter reads it when present (see evals/helpers.py).
    result: str
    steps_taken: list[str]


supervisor_agent: Agent[SharedDeps, SupervisorOutput] = Agent(
    settings.model,
    name=LABEL,
    output_type=SupervisorOutput,
    deps_type=SharedDeps,
    # Fail fast when the provider filters a response, instead of retrying a
    # refused request or returning partial text.
    capabilities=[RaiseContentFilterError()],
    instructions="""You coordinate two workers to answer the user's request.

    - delegate_to_analyst: researches a question and returns key findings.
    - delegate_to_writer: turns material into prose. It knows only what you pass it, so
      include the findings and any instructions about length or audience in `task`.

    For a request that needs research and a written answer, call the analyst first, then the
    writer with the analyst's findings. If one worker is enough, call only that one. Return
    the final text in `result`, and in `steps_taken` list each worker you called, in order.
    """,
)


# --- Supervisor tools that delegate to workers ---
@supervisor_agent.tool
async def delegate_to_analyst(ctx: RunContext[SharedDeps], task: str) -> str:
    """Ask the analyst to research a question or topic.

    Args:
        task: The question or topic to analyze.

    Returns:
        The analyst's key findings, one per line.
    """
    logger.info("Delegating to analyst", extra={"task": task})
    # usage=ctx.usage makes the worker's spend count against the supervisor
    # run's shared budget — the standard pydantic-ai delegation pattern.
    result = await analyst_agent.run(
        task, deps=ctx.deps, usage=ctx.usage, usage_limits=USAGE_LIMITS
    )
    return "\n".join(f"- {point}" for point in result.output.points)


@supervisor_agent.tool
async def delegate_to_writer(ctx: RunContext[SharedDeps], task: str) -> str:
    """Ask the writer to turn material into prose.

    Args:
        task: The material to write from, plus any instructions about length or audience.

    Returns:
        The written text.
    """
    logger.info("Delegating to writer", extra={"task": task})
    result = await writer_agent.run(task, deps=ctx.deps, usage=ctx.usage, usage_limits=USAGE_LIMITS)
    return result.output.text


async def run_supervisor(
    user_input: str, deps: SharedDeps | None = None
) -> RunResult[SupervisorOutput]:
    """Run the supervisor agent to coordinate workers on a task.

    Args:
        user_input: The user's message or task description.
        deps: Runtime dependencies. Created with defaults if not provided.

    Returns:
        A RunResult with one step, the supervisor's. The workers it delegated to ran inside
        that step (their calls are in `.all_messages()`), and `.usage` includes their spend.
    """
    if deps is None:
        deps = SharedDeps()
    logger.info("Running supervisor agent", extra={"user_input": user_input})
    flow = Flow(USAGE_LIMITS)
    result = await flow.run(supervisor_agent, user_input, deps=deps)
    return flow.finish(result.output)


if __name__ == "__main__":
    import asyncio

    configure_logging()
    result = asyncio.run(
        run_supervisor(
            "Research the pros and cons of remote work, then write two sentences about it for a manager."
        )
    )
    print(result.output)
title = "Supervisor / workers"
pattern = "supervisor"
summary = "A supervisor decides which specialized workers to call, and in what order, then synthesizes their results."
smoke_input = "Research the pros and cons of remote work, then write two sentences about it for a manager."
# A real model must delegate to both workers for this input (checked by the release check).
expected_tools = ["delegate_to_analyst", "delegate_to_writer"]

# The offline smoke test opts in to both delegation tools, so both workers really run.
[smoke.supervisor_agent]
call_tools = ["delegate_to_analyst", "delegate_to_writer"]

[entrypoint]
deps = "SharedDeps"
run = "run_supervisor"
"""The supervisor delegates to real workers, passes their results along, and shares one budget."""

import pytest
from pydantic_ai.exceptions import UsageLimitExceeded
from pydantic_ai.messages import ModelResponse, ToolCallPart, ToolReturnPart, UserPromptPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
from pydantic_ai.usage import UsageLimits

from examples.supervisor import agent as supervisor
from examples.supervisor.agent import (
    analyst_agent,
    run_supervisor,
    supervisor_agent,
    writer_agent,
)


def prompt_of(messages) -> str:
    return "\n".join(
        str(p.content) for m in messages for p in m.parts if isinstance(p, UserPromptPart)
    )


def final(info: AgentInfo, fields: dict) -> ModelResponse:
    return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, fields)])


def worker(fields: dict, seen: list[str] | None = None):
    """A worker that records the prompt it was given and returns `fields`."""

    def model_fn(messages, info: AgentInfo) -> ModelResponse:
        if seen is not None:
            seen.append(prompt_of(messages))
        return final(info, fields)

    return FunctionModel(model_fn)


def supervisor_that(*plan: tuple[str, str] | dict):
    """A supervisor that makes the delegation calls in `plan`, then returns the final dict."""
    remaining = list(plan)

    def model_fn(messages, info: AgentInfo) -> ModelResponse:
        step = remaining.pop(0)
        if isinstance(step, dict):
            return final(info, step)
        tool, task = step
        return ModelResponse(parts=[ToolCallPart(tool, {"task": task})])

    return FunctionModel(model_fn)


def tool_outputs(result) -> dict[str, str]:
    return {
        p.tool_name: str(p.content)
        for m in result.all_messages()
        for p in m.parts
        if isinstance(p, ToolReturnPart) and p.tool_name != "final_result"
    }


async def test_research_then_write_hands_the_findings_to_the_writer():
    writer_saw: list[str] = []
    plan = supervisor_that(
        ("delegate_to_analyst", "monorepos"),
        ("delegate_to_writer", "Write one sentence from: - one repo\n- shared tooling"),
        {"result": "Monorepos share tooling.", "steps_taken": ["analyst", "writer"]},
    )
    with (
        supervisor_agent.override(model=plan),
        analyst_agent.override(model=worker({"points": ["one repo", "shared tooling"]})),
        writer_agent.override(model=worker({"text": "Monorepos share tooling."}, writer_saw)),
    ):
        result = await run_supervisor("Explain monorepos")

    assert result.output.steps_taken == ["analyst", "writer"]
    outputs = tool_outputs(result)
    assert outputs["delegate_to_analyst"] == "- one repo\n- shared tooling"  # one point per line
    assert outputs["delegate_to_writer"] == "Monorepos share tooling."
    assert "shared tooling" in writer_saw[0]  # the supervisor passed the findings on


async def test_the_supervisor_may_call_just_one_worker():
    plan = supervisor_that(
        ("delegate_to_writer", "Say hello"),
        {"result": "Hello", "steps_taken": ["writer"]},
    )
    with (
        supervisor_agent.override(model=plan),
        writer_agent.override(model=worker({"text": "Hello"})),
    ):
        result = await run_supervisor("Say hello")
    assert set(tool_outputs(result)) == {"delegate_to_writer"}


async def test_the_result_has_one_step_and_the_workers_spend_is_in_the_total():
    plan = supervisor_that(
        ("delegate_to_analyst", "x"), {"result": "r", "steps_taken": ["analyst"]}
    )
    with (
        supervisor_agent.override(model=plan),
        analyst_agent.override(model=worker({"points": ["p"]})),
    ):
        result = await run_supervisor("x")

    assert [step.agent for step in result.steps] == ["supervisor"]
    # Two supervisor requests (the delegation, then the answer) plus the analyst's one.
    assert result.usage.requests == 3


async def test_the_workers_share_the_supervisors_budget(monkeypatch):
    """request_limit=2 allows the supervisor's two requests but not the worker's third."""
    monkeypatch.setattr(supervisor, "USAGE_LIMITS", UsageLimits(request_limit=2))
    plan = supervisor_that(
        ("delegate_to_analyst", "x"), {"result": "r", "steps_taken": ["analyst"]}
    )
    with (
        supervisor_agent.override(model=plan),
        analyst_agent.override(model=worker({"points": ["p"]})),
        pytest.raises(UsageLimitExceeded),
    ):
        await run_supervisor("x")
"""Live check: the supervisor delegates to real workers and synthesizes. Run with `pytest -m eval`."""

import pytest

from evals.trace import traced_run
from examples.live_support import assert_every_agent_ran, run_as_script
from examples.supervisor import agent as module

pytestmark = pytest.mark.eval


@pytest.fixture(scope="module")
async def research_and_write():
    return await traced_run(
        module.run_supervisor,
        "Research the pros and cons of remote work, then write two sentences about it for a manager.",
    )


@pytest.fixture(scope="module")
async def write_only():
    return await traced_run(module.run_supervisor, "Say hello in French.")


async def test_a_research_and_write_request_uses_both_workers_in_order(research_and_write):
    traced = research_and_write
    assert traced.tools_called == ["delegate_to_analyst", "delegate_to_writer"]
    assert traced.agents_ran == {"supervisor", "supervisor.analyst", "supervisor.writer"}

    output = traced.result.output
    assert len(output.result.split()) >= 10
    # steps_taken is the model's own account; it should name both workers, analyst first.
    named = [s.lower() for s in output.steps_taken]
    assert any("analyst" in s for s in named) and any("writer" in s for s in named)
    assert min(i for i, s in enumerate(named) if "analyst" in s) < max(
        i for i, s in enumerate(named) if "writer" in s
    )
    # The workers' requests count against the one shared budget.
    assert traced.result.usage.requests >= 5
    assert [step.agent for step in traced.result.steps] == ["supervisor"]


async def test_a_simple_request_is_answered_without_needing_every_worker(write_only):
    # Whether the model uses the writer or answers directly is its call; either way, it answers.
    assert any(w in write_only.result.output.result.lower() for w in ("bonjour", "salut"))


async def test_every_agent_ran(research_and_write, write_only):
    assert_every_agent_ran(module, research_and_write.agents_ran | write_only.agents_ran)


async def test_the_demo_script_runs():
    assert "result=" in await run_as_script("examples.supervisor.agent")

Recorded run · gemini-3.1-flash-lite · 1 step · $0.0008

Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 2,312 tokens · $0.0008 · 8.4 s. Model output varies between runs. Regenerate with uv run python scripts/record_example.py supervisor.

Input

Research the pros and cons of remote work, then write two sentences about it for a manager.

Steps

1. supervisor 1,877 tokens · $0.0008

Prompt

Research the pros and cons of remote work, then write two sentences about it for a manager.

What happened - called delegate_to_analyst({"task": "Research the pros and cons of remote work."}) - delegate_to_analyst returned: - Remote work enhances employee flexibility, reduces commuting time, and lowers overhead costs for organizations. - Workers often experience improved work-life balance and increased productivity due t … - called delegate_to_writer({"task": "Using these findings:\n- Remote work enhances employee flexibility, reduces commuting time, and lowers overhead costs for organizations.\n- Workers often experience improved work-life balanc …) - delegate_to_writer returned: Remote work benefits managers by lowering overhead costs and increasing employee productivity through fewer distractions. However, these advantages are balanced by challenges such as weakened team coh …

Output

{
  "result": "Remote work benefits managers by lowering overhead costs and increasing employee productivity through fewer distractions. However, these advantages are balanced by challenges such as weakened team cohesion, reduced spontaneous collaboration, and a higher risk of employee burnout.",
  "steps_taken": [
    "delegate_to_analyst",
    "delegate_to_writer"
  ]
}

Result

run_supervisor(...).output

{
  "result": "Remote work benefits managers by lowering overhead costs and increasing employee productivity through fewer distractions. However, these advantages are balanced by challenges such as weakened team cohesion, reduced spontaneous collaboration, and a higher risk of employee burnout.",
  "steps_taken": [
    "delegate_to_analyst",
    "delegate_to_writer"
  ]
}