Skip to content

Router

A cheap classifier picks a category, and plain code sends the input to the right specialist.

A router handles different kinds of input differently. A small classifier agent reads the input and returns one category from a fixed list (its output type allows nothing else). Plain code then looks the category up in a dictionary and hands the input to that category's specialist agent. The model decides only what kind of thing this is; what happens next is ordinary Python, so it is predictable, easy to test and cheap, and there is no agent loop to bound.

Use it when

  • Inputs fall into a known set of categories.
  • Each category is best handled by its own agent, with its own instructions, tools or tone.
  • You want routing that is predictable, testable and cheap.

Look elsewhere when

  • The categories aren't known in advance, or one request needs several specialists: supervisor.
  • Every input gets the same handling: single.
classifier → category → SPECIALISTS[category] → answer

What it shows

  • Classification as a Literal output type: the model can only return a known category
  • Dispatch in code (SPECIALISTS[category]), each specialist with its own instructions
  • One RunUsage shared across the classifier and the specialist, so USAGE_LIMITS bounds the whole route
  • Trace labels <name>.classifier, <name>.billing, … so each hop is identifiable

See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.

uv run python scripts/add_agent.py router --name support

run_router returns a RunResult: .output is the answer and its category, and .steps records the route taken (the classifier, then that category's specialist).

Adapt it by changing Category and SPECIALISTS. examples/router/test_example.py tests the dispatch and the shared budget.

Source

All of it is in examples/router/.

"""Router: a cheap classifier picks a specialist, and plain code does the dispatch.

Use this pattern when:
- Inputs fall into a known set of categories, each best handled by its own agent
- You want routing to be predictable, testable and cheap — not another LLM decision loop

How it differs from `supervisor`: there, the LLM decides which worker to call by using
delegation tools. Here the model only *classifies*; a dictionary in your code maps the
category to a specialist. Fewer moving parts, and the routing is plain Python you can test.

    classifier → category → SPECIALISTS[category] → answer
"""

from __future__ import annotations

from dataclasses import dataclass
from typing import Literal

from pydantic import BaseModel
from pydantic_ai import Agent
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits

from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.runs import Flow, RunResult

logger = get_logger(__name__)
LABEL = agent_label(__name__)  # names this agent's run spans in Logfire traces

# One budget for the whole route: the classifier and the specialist run in one Flow
# (see run_router), so these limits bound the pair, not each call separately.
USAGE_LIMITS = UsageLimits(
    request_limit=10, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)

Category = Literal["billing", "technical", "general"]


@dataclass
class RouterDeps:
    """Runtime dependencies shared by the classifier and the specialists."""

    pass


# --- Classifier ---
class Classification(BaseModel):
    category: Category


router_agent: Agent[RouterDeps, Classification] = Agent(
    settings.model,
    name=f"{LABEL}.classifier",
    output_type=Classification,
    deps_type=RouterDeps,
    capabilities=[RaiseContentFilterError()],
    instructions="""Classify the support message into exactly one category:

    - billing: invoices, payments, refunds, plans and pricing
    - technical: errors, bugs, outages, how-to questions about the product
    - general: anything else
    """,
)


# --- Specialists ---
class Answer(BaseModel):
    result: str


def _specialist(category: Category, instructions: str) -> Agent[RouterDeps, Answer]:
    return Agent(
        settings.model,
        name=f"{LABEL}.{category}",
        output_type=Answer,
        deps_type=RouterDeps,
        capabilities=[RaiseContentFilterError()],
        instructions=instructions,
    )


# One specialist per category. Each can have its own instructions, tools, even model.
billing_agent = _specialist("billing", "You are a billing support specialist. Be precise.")
technical_agent = _specialist("technical", "You are a technical support engineer. Be concrete.")
general_agent = _specialist("general", "You are a friendly support agent. Keep answers short.")

SPECIALISTS: dict[Category, Agent[RouterDeps, Answer]] = {
    "billing": billing_agent,
    "technical": technical_agent,
    "general": general_agent,
}


class RouterOutput(BaseModel):
    result: str
    category: Category


async def run_router(user_input: str, deps: RouterDeps | None = None) -> RunResult[RouterOutput]:
    """Classify `user_input`, then answer it with the matching specialist.

    Returns:
        A RunResult: `.output` is the RouterOutput; `.steps` holds the classifier step and then
        the specialist step.
    """
    if deps is None:
        deps = RouterDeps()
    flow = Flow(USAGE_LIMITS)  # one shared budget, so USAGE_LIMITS bounds the whole route

    classified = await flow.run(router_agent, user_input, deps=deps)
    category = classified.output.category
    logger.info("Routed", extra={"category": category})

    # The routing decision is a dictionary lookup, not an LLM call.
    answered = await flow.run(SPECIALISTS[category], user_input, deps=deps)
    return flow.finish(RouterOutput(result=answered.output.result, category=category))


if __name__ == "__main__":
    import asyncio

    configure_logging()
    print(asyncio.run(run_router("I was charged twice for my subscription this month.")).output)
title = "Router"
pattern = "router"
summary = "A classifier picks a category and plain code dispatches to a specialist agent; routing is a lookup, not an LLM loop."
smoke_input = "I was charged twice for my subscription this month."

[entrypoint]
deps = "RouterDeps"
run = "run_router"
"""Routing is a dictionary lookup: each category reaches its own specialist."""

import pytest
from pydantic import ValidationError
from pydantic_ai.exceptions import UsageLimitExceeded
from pydantic_ai.messages import ModelResponse, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
from pydantic_ai.usage import UsageLimits

from examples.router import agent as router
from examples.router.agent import (
    SPECIALISTS,
    Classification,
    router_agent,
    run_router,
)


def returns(**fields):
    def model_fn(messages, info: AgentInfo) -> ModelResponse:
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, fields)])

    return FunctionModel(model_fn)


@pytest.mark.parametrize("category", sorted(SPECIALISTS))
async def test_each_category_reaches_its_own_specialist(category):
    # Every specialist answers with its own name, so a misroute shows in the result.
    overrides = [SPECIALISTS[c].override(model=returns(result=f"{c} answer")) for c in SPECIALISTS]
    with router_agent.override(model=returns(category=category)):
        for override in overrides:
            override.__enter__()
        try:
            result = await run_router("anything")
            output = result.output
        finally:
            for override in reversed(overrides):
                override.__exit__(None, None, None)

    assert output.category == category
    assert output.result == f"{category} answer"
    # The result records the route taken: the classifier, then that category's specialist.
    assert [step.agent for step in result.steps] == ["router.classifier", f"router.{category}"]
    assert result.usage.requests == 2


def test_the_classifier_can_only_return_a_known_category():
    with pytest.raises(ValidationError):
        Classification(category="astrology")


async def test_one_budget_covers_the_classifier_and_the_specialist(monkeypatch):
    """USAGE_LIMITS bounds the whole route: two requests exceed request_limit=1."""
    monkeypatch.setattr(router, "USAGE_LIMITS", UsageLimits(request_limit=1))
    with (
        router_agent.override(model=returns(category="general")),
        SPECIALISTS["general"].override(model=returns(result="hi")),
        pytest.raises(UsageLimitExceeded),
    ):
        await run_router("hello")
"""Live check: each message reaches the right specialist, and every specialist runs. `-m eval`."""

import pytest

from evals.trace import traced_run
from examples.live_support import assert_every_agent_ran, run_as_script
from examples.router import agent as module

pytestmark = pytest.mark.eval

CASES = {
    "billing": "I was charged twice for my subscription this month.",
    "technical": "The app crashes with a 500 error every time I upload a file.",
    "general": "Do you have an office in Berlin?",
}


@pytest.fixture(scope="module")
async def routed():
    return {
        category: await traced_run(module.run_router, message)
        for category, message in CASES.items()
    }


@pytest.mark.parametrize("category", sorted(CASES))
async def test_a_message_is_classified_and_answered_by_its_specialist(routed, category):
    traced = routed[category]
    output = traced.result.output

    assert output.category == category
    assert output.result.strip()
    assert [step.agent for step in traced.result.steps] == [
        "router.classifier",
        f"router.{category}",
    ]
    assert traced.agents_ran == {"router.classifier", f"router.{category}"}
    assert traced.result.usage.requests >= 2  # one per step, plus any output retries


async def test_every_specialist_ran_for_real(routed):
    assert_every_agent_ran(module, set().union(*(t.agents_ran for t in routed.values())))


async def test_the_demo_script_runs():
    assert "category=" in await run_as_script("examples.router.agent")

Recorded run · gemini-3.1-flash-lite · 2 steps · $0.0003

Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 2 steps · 334 tokens · $0.0003 · 1.8 s. Model output varies between runs. Regenerate with uv run python scripts/record_example.py router.

Input

I was charged twice for my subscription this month.

Steps

1. router.classifier 141 tokens · $0.0001

Prompt

I was charged twice for my subscription this month.

Output

{
  "category": "billing"
}

2. router.billing 193 tokens · $0.0002

Prompt

I was charged twice for my subscription this month.

Output

{
  "result": "I’m sorry to hear that you were charged twice. I can certainly help you look into this.\n\nTo get started, please provide the following information:\n\n1. The email address associated with your subscription account.\n2. The dates and amounts of the two separate charges.\n3. If you have them, the transaction IDs or order numbers for both charges.\n\nOnce I have those details, I will investigate the billing records and process a refund for the duplicate charge if confirmed. Please do not share any full credit card numbers or passwords here."
}

Result

run_router(...).output

{
  "result": "I’m sorry to hear that you were charged twice. I can certainly help you look into this.\n\nTo get started, please provide the following information:\n\n1. The email address associated with your subscription account.\n2. The dates and amounts of the two separate charges.\n3. If you have them, the transaction IDs or order numbers for both charges.\n\nOnce I have those details, I will investigate the billing records and process a refund for the duplicate charge if confirmed. Please do not share any full credit card numbers or passwords here.",
  "category": "billing"
}