Evaluator–optimizer¶
A generator drafts, a critic reviews against criteria, and they loop until the draft passes.
Two agents work in a loop. A generator writes a draft. A critic reviews it against criteria you wrote and either accepts it or returns specific feedback, and the feedback goes back to the generator for another try. The loop ends when the critic accepts or a round cap is reached; hitting the cap returns the last draft marked as not accepted, so the caller decides what to do. It trades extra model calls for a better answer, and how good the answer gets depends on how well you can state what good looks like.
Use it when
- You can state what "good" looks like as criteria a second model can check.
- A first attempt is usually close but benefits from targeted revision.
- A few extra model calls are worth a better answer.
Look elsewhere when
- "Good" can't be written down, or is better checked in code: a validator (
extraction) is cheaper. - A first attempt is usually fine:
single.
generator → draft → critic ─ accepted ─→ done
▲ │
└── feedback ────┘ (at most MAX_ITERATIONS rounds)
What it shows
- Two agents with a structured
Critique(accepted,feedback) closing the loop - A round cap (
MAX_ITERATIONS) on top ofUSAGE_LIMITS; hitting it returns the last draft withaccepted=Falseinstead of looping or raising, so the caller decides - Two prompt files,
evaluator_optimizer_generator.txtandevaluator_optimizer_critic.txt(add_agent.pyrenames both to your agent's name) - One
RunUsageshared across rounds; trace labels<name>.generatorand<name>.critic
See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.
run_evaluator_optimizer returns a RunResult: .output carries accepted and iterations,
and .steps holds every generate and critique round in order.
Adapt it by rewriting the critic's criteria, which are the whole point.
examples/evaluator_optimizer/test_example.py tests the loop, the feedback and the cap.
Source¶
All of it is in examples/evaluator_optimizer/.
"""Evaluator–optimizer: a generator drafts, a critic reviews, and they loop until it passes.
Use this pattern when:
- You can state what "good" looks like as criteria a second model can check
- A first attempt is usually close but benefits from targeted revision
- The cost of a few extra model calls is worth a better final answer
generator → draft → critic ─ accepted ─→ done
▲ │
└── feedback ────┘ (at most MAX_ITERATIONS rounds)
The loop is bounded twice: by an explicit round cap, and by `USAGE_LIMITS`.
"""
from __future__ import annotations
from dataclasses import dataclass
from pydantic import BaseModel
from pydantic_ai import Agent
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits
from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult
logger = get_logger(__name__)
LABEL = agent_label(__name__) # names this agent's run spans in Logfire traces
# Each round costs two requests (generate + critique). Size request_limit for the cap below.
USAGE_LIMITS = UsageLimits(
request_limit=12, total_tokens_limit=150_000, cost_limit=settings.cost_limit
)
# Hard cap on revise rounds. Without one, a critic that is never satisfied loops until
# USAGE_LIMITS trips; with one, you decide what "good enough" means.
MAX_ITERATIONS = 3
@dataclass
class EvaluatorOptimizerDeps:
"""Runtime dependencies shared by the generator and the critic."""
pass
class Draft(BaseModel):
result: str
generator_agent: Agent[EvaluatorOptimizerDeps, Draft] = Agent(
settings.model,
name=f"{LABEL}.generator",
output_type=Draft,
deps_type=EvaluatorOptimizerDeps,
capabilities=[RaiseContentFilterError()],
instructions=load_prompt("evaluator_optimizer_generator"),
)
class Critique(BaseModel):
accepted: bool
feedback: str
critic_agent: Agent[EvaluatorOptimizerDeps, Critique] = Agent(
settings.model,
name=f"{LABEL}.critic",
output_type=Critique,
deps_type=EvaluatorOptimizerDeps,
capabilities=[RaiseContentFilterError()],
instructions=load_prompt("evaluator_optimizer_critic"),
)
class EvaluatorOptimizerOutput(BaseModel):
result: str
# False means the round cap was hit: `result` is the best draft so far, not a passing one.
accepted: bool
iterations: int
async def run_evaluator_optimizer(
user_input: str, deps: EvaluatorOptimizerDeps | None = None
) -> RunResult[EvaluatorOptimizerOutput]:
"""Draft a description of `user_input`, revising until the critic accepts or the cap is hit.
Hitting the cap returns the last draft with `accepted=False` rather than raising; the
caller decides whether that is good enough. `.steps` on the result holds every generate and
critique round, in order.
"""
if deps is None:
deps = EvaluatorOptimizerDeps()
flow = Flow(USAGE_LIMITS) # one shared budget, so USAGE_LIMITS bounds every round
prompt = f"Item: {user_input}"
draft = ""
for iteration in range(1, MAX_ITERATIONS + 1):
draft = (await flow.run(generator_agent, prompt, deps=deps)).output.result
review = f"Item: {user_input}\n\nDescription:\n{draft}"
critique = (await flow.run(critic_agent, review, deps=deps)).output
logger.info("Reviewed", extra={"iteration": iteration, "accepted": critique.accepted})
if critique.accepted:
return flow.finish(
EvaluatorOptimizerOutput(result=draft, accepted=True, iterations=iteration)
)
# Feed the critic's feedback back in, along with the draft it was about.
prompt = (
f"Item: {user_input}\n\nPrevious attempt:\n{draft}\n\nFeedback to address:\n"
f"{critique.feedback}"
)
return flow.finish(
EvaluatorOptimizerOutput(result=draft, accepted=False, iterations=MAX_ITERATIONS)
)
if __name__ == "__main__":
import asyncio
configure_logging()
print(asyncio.run(run_evaluator_optimizer("A stainless steel water bottle")).output)
You review a product description against these criteria:
- It is two or three sentences long.
- It says what the item is and gives one concrete benefit.
- It makes no factual claims that the item's name does not support: no invented materials,
performance figures, certifications or statistics. Ordinary descriptive wording is fine.
Accept it only if it meets all of them. Otherwise do not accept it, and give specific
feedback the writer can act on.
title = "Evaluator–optimizer"
pattern = "evaluator_optimizer"
summary = "A generator drafts and a critic reviews against criteria; they loop until it passes or a round cap is hit."
smoke_input = "A stainless steel water bottle"
[entrypoint]
deps = "EvaluatorOptimizerDeps"
run = "run_evaluator_optimizer"
"""The loop stops when the critic accepts, feeds feedback back in, and is capped."""
from pydantic_ai.messages import ModelResponse, ToolCallPart, UserPromptPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
from examples.evaluator_optimizer.agent import (
MAX_ITERATIONS,
critic_agent,
generator_agent,
run_evaluator_optimizer,
)
def prompt_of(messages) -> str:
return "\n".join(str(p.content) for p in messages[-1].parts if isinstance(p, UserPromptPart))
def generator(seen: list[str]):
"""Drafts `draft 1`, `draft 2`, ... and records each prompt."""
def model_fn(messages, info: AgentInfo) -> ModelResponse:
seen.append(prompt_of(messages))
fields = {"result": f"draft {len(seen)}"}
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, fields)])
return FunctionModel(model_fn)
def critic(verdicts: list[bool]):
"""Returns the given accept/reject verdicts in order; every rejection carries feedback."""
remaining = list(verdicts)
def model_fn(messages, info: AgentInfo) -> ModelResponse:
fields = {"accepted": remaining.pop(0), "feedback": "make it shorter"}
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, fields)])
return FunctionModel(model_fn)
async def test_a_draft_the_critic_accepts_is_returned_after_one_round():
prompts: list[str] = []
with (
generator_agent.override(model=generator(prompts)),
critic_agent.override(model=critic([True])),
):
result = await run_evaluator_optimizer("a bottle")
output = result.output
assert (output.result, output.accepted, output.iterations) == ("draft 1", True, 1)
assert len(prompts) == 1
async def test_feedback_and_the_previous_draft_go_back_to_the_generator():
prompts: list[str] = []
with (
generator_agent.override(model=generator(prompts)),
critic_agent.override(model=critic([False, True])),
):
result = await run_evaluator_optimizer("a bottle")
output = result.output
assert (output.result, output.accepted, output.iterations) == ("draft 2", True, 2)
assert "make it shorter" in prompts[1]
assert [step.agent for step in result.steps] == [
"evaluator_optimizer.generator",
"evaluator_optimizer.critic",
"evaluator_optimizer.generator",
"evaluator_optimizer.critic",
]
assert "draft 1" in prompts[1]
async def test_a_critic_that_is_never_satisfied_is_capped_not_looped():
prompts: list[str] = []
with (
generator_agent.override(model=generator(prompts)),
critic_agent.override(model=critic([False] * MAX_ITERATIONS)),
):
result = await run_evaluator_optimizer("a bottle")
output = result.output
assert output.accepted is False
assert output.iterations == MAX_ITERATIONS
assert output.result == f"draft {MAX_ITERATIONS}" # the best attempt so far, not an error
"""Live check: the generate/critique loop's bookkeeping holds on real model output. `-m eval`.
How many rounds the real critic takes varies from run to run, so these assert the loop's
invariants for whichever path occurs rather than a particular verdict.
"""
import pytest
from evals.trace import traced_run
from examples.evaluator_optimizer import agent as module
from examples.live_support import assert_every_agent_ran, run_as_script
pytestmark = pytest.mark.eval
ITEMS = ["A stainless steel water bottle", "A wireless mechanical keyboard"]
@pytest.fixture(scope="module")
async def runs():
return [await traced_run(module.run_evaluator_optimizer, item) for item in ITEMS]
async def test_the_loop_stops_for_a_real_reason_and_its_bookkeeping_is_consistent(runs):
for traced in runs:
output = traced.result.output
steps = traced.result.steps
assert 1 <= output.iterations <= module.MAX_ITERATIONS
# It stopped because the critic accepted, or because it hit the cap — nothing else.
assert output.accepted or output.iterations == module.MAX_ITERATIONS
# One generate step and one critique step per round, alternating.
assert [s.agent for s in steps] == [
"evaluator_optimizer.generator",
"evaluator_optimizer.critic",
] * output.iterations
# The returned text is the last draft, and `accepted` is the last verdict.
assert output.result == steps[-2].result.output.result
assert output.accepted == steps[-1].result.output.accepted
assert output.result.strip()
async def test_feedback_from_a_rejection_reaches_the_next_draft(runs):
revised = [t for t in runs if t.result.output.iterations > 1]
if not revised:
pytest.skip("the real critic accepted every first draft, so no revision happened")
steps = revised[0].result.steps
rejection = steps[1].result.output
assert rejection.accepted is False and rejection.feedback.strip()
second_prompt = next(
str(part.content)
for message in steps[2].result.all_messages()
for part in message.parts
if type(part).__name__ == "UserPromptPart"
)
assert rejection.feedback in second_prompt and steps[0].result.output.result in second_prompt
async def test_both_agents_ran(runs):
assert_every_agent_ran(module, set().union(*(t.agents_ran for t in runs)))
async def test_the_demo_script_runs():
assert "iterations=" in await run_as_script("examples.evaluator_optimizer.agent")
Recorded run · gemini-3.1-flash-lite · 6 steps · $0.0010
Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 6 steps · 1,731 tokens · $0.0010 · 6.8 s.
Model output varies between runs. Regenerate with uv run python scripts/record_example.py evaluator_optimizer.
Input
A stainless steel water bottle
Steps
1. evaluator_optimizer.generator
185 tokens · $0.0002
Prompt
Item: A stainless steel water bottle
Output
{
"result": "Stay hydrated on the go with this durable stainless steel water bottle. Crafted from high-quality, food-grade steel, it features a sleek, leak-proof design that fits perfectly in your cup holder—making it the ideal companion for your daily commute, gym sessions, or outdoor adventures. Keep your drinks at the perfect temperature for hours while reducing single-use plastic waste."
}
2. evaluator_optimizer.critic
326 tokens · $0.0002
Prompt
Item: A stainless steel water bottle
Description: Stay hydrated on the go with this durable stainless steel water bottle. Crafted from high-quality, food-grade steel, it features a sleek, leak-proof design that fits perfectly in your cup holder—making it the ideal companion for your daily commute, gym sessions, or outdoor adventures. Keep your drinks at the perfect temperature for hours while reducing single-use plastic waste.
Output
{
"accepted": false,
"feedback": "The description exceeds the three-sentence limit. Additionally, it makes factual claims (\"high-quality,\" \"food-grade,\" \"leak-proof,\" and \"keep your drinks at the perfect temperature for hours\") that are not supported by the product name. Please remove these claims and condense the text to two or three sentences."
}
3. evaluator_optimizer.generator
309 tokens · $0.0002
Prompt
Item: A stainless steel water bottle
Previous attempt: Stay hydrated on the go with this durable stainless steel water bottle. Crafted from high-quality, food-grade steel, it features a sleek, leak-proof design that fits perfectly in your cup holder—making it the ideal companion for your daily commute, gym sessions, or outdoor adventures. Keep your drinks at the perfect temperature for hours while reducing single-use plastic waste.
Feedback to address: The description exceeds the three-sentenc …
Output
{
"result": "Stay hydrated anywhere with this durable stainless steel water bottle. Its versatile design fits most cup holders, making it a convenient choice for your commute, gym sessions, or outdoor adventures. This reusable bottle helps you reduce reliance on single-use plastics."
}
4. evaluator_optimizer.critic
363 tokens · $0.0003
Prompt
Item: A stainless steel water bottle
Description: Stay hydrated anywhere with this durable stainless steel water bottle. Its versatile design fits most cup holders, making it a convenient choice for your commute, gym sessions, or outdoor adventures. This reusable bottle helps you reduce reliance on single-use plastics.
Output
{
"accepted": false,
"feedback": "The description is three sentences long, which is fine, but it fails the length constraint of the instructions. The instructions require the description to be \"two or three sentences long,\" however, the description provided is actually three sentences. Wait—it *is* three sentences. Let me re-read. Ah, it meets the length requirement. However, it fails the \"it says what the item is and gives one concrete benefit\" requirement. You have included three benefits (fits cup holders, convenient for commute/gym/etc., and helps reduce plastic use) rather than \"one concrete benefit.\" Please trim the description to include only one benefit."
}
5. evaluator_optimizer.generator
327 tokens · $0.0001
Prompt
Item: A stainless steel water bottle
Previous attempt: Stay hydrated anywhere with this durable stainless steel water bottle. Its versatile design fits most cup holders, making it a convenient choice for your commute, gym sessions, or outdoor adventures. This reusable bottle helps you reduce reliance on single-use plastics.
Feedback to address: The description is three sentences long, which is fine, but it fails the length constraint of the instructions. The instructions require the descriptio …
Output
{
"result": "Stay hydrated anywhere with this durable stainless steel water bottle. Its versatile design fits most cup holders, making it a convenient choice for your daily commute."
}
6. evaluator_optimizer.critic
221 tokens · $0.0001
Prompt
Item: A stainless steel water bottle
Description: Stay hydrated anywhere with this durable stainless steel water bottle. Its versatile design fits most cup holders, making it a convenient choice for your daily commute.
Output
Result
run_evaluator_optimizer(...).output