Structured extraction¶
Turn unstructured text into a validated schema, and send the model back to fix what doesn't validate.
Extraction asks the model to fill in a typed record from free text: an email, a ticket, a document. Pydantic checks the shape, and an output validator you write checks the content. When either check fails, the reason goes back to the model in plain English and it tries again, up to a retry budget you set; when the budget runs out, the run raises instead of returning something half-valid. The result is data you can pass downstream without checking it again.
Use it when
- The input is free text (emails, tickets, documents) and you need typed fields out of it.
- Wrong output should be caught and corrected, not passed on.
- You can say, in code, what a valid answer looks like.
Look elsewhere when
- The output is prose for a person to read:
singleis enough. - The schema is large or deeply nested: split the job (see
pipeline), because nesting the data doesn't need makes smaller models fail.
What it shows
- A flat output schema (
Contact); nesting the data doesn't need makes smaller models fail - An output validator that raises
ModelRetrywith a plain-English reason, so the model gets to fix its own answer - An explicit retry budget (
retries={"output": 2}); when it runs out the run raisesUnexpectedModelBehaviorinstead of returning something half-valid [smoke.extraction_agent]inexample.toml:TestModelgenerates junk that a real validator rejects, so the offline tests are told what to return
See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.
Adapt it by replacing Contact with your schema and check_contact with your own checks.
Source¶
All of it is in examples/extraction/.
"""Structured extraction: turn unstructured text into a validated schema.
Use this pattern when:
- The input is free text (emails, tickets, documents) and you need typed fields out
- Bad output should be caught and corrected, not passed downstream
How it works:
1. `Contact` is the schema the model must fill in; keep it as flat as the data allows
2. An output validator checks what the model returned and raises `ModelRetry` with a
plain-English reason, so the model gets a chance to fix its own answer
3. The retry budget is explicit; when it runs out the run fails loudly instead of
returning something half-valid
"""
from __future__ import annotations
from dataclasses import dataclass
from pydantic import BaseModel
from pydantic_ai import Agent, ModelRetry, RunContext
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.usage import UsageLimits
from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult
logger = get_logger(__name__)
LABEL = agent_label(__name__) # names this agent's run spans in Logfire traces
# Guardrail against runaway agentic loops (see examples/single for the details).
USAGE_LIMITS = UsageLimits(
request_limit=10, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)
# --- Output type ---
# Flat on purpose: nesting the data doesn't need makes small models fail validation.
class Contact(BaseModel):
name: str
email: str | None = None
phone: str | None = None
company: str | None = None
# --- Dependencies ---
@dataclass
class ExtractionDeps:
"""Runtime dependencies injected into the extraction agent."""
pass
# --- Agent definition ---
extraction_agent: Agent[ExtractionDeps, Contact] = Agent(
settings.model,
name=LABEL,
output_type=Contact,
deps_type=ExtractionDeps,
# How many times the model may be sent back to fix a rejected answer.
retries={"output": 2},
capabilities=[RaiseContentFilterError()],
instructions=load_prompt("extraction"),
)
# --- Validation ---
@extraction_agent.output_validator
def check_contact(ctx: RunContext[ExtractionDeps], contact: Contact) -> Contact:
"""Reject answers the model can plausibly fix; the message is sent back to it.
ModelRetry is for errors the model can correct by changing its answer. Anything
that is a bug in this code should raise normally instead.
"""
if contact.email is not None and "@" not in contact.email:
raise ModelRetry(
f"'{contact.email}' is not an email address. Copy the address exactly as it "
"appears in the text, or leave email empty if there isn't one."
)
if contact.email is None and contact.phone is None:
raise ModelRetry("A contact needs an email address or a phone number; neither was found.")
return contact
async def run_extraction(user_input: str, deps: ExtractionDeps | None = None) -> RunResult[Contact]:
"""Extract a contact from `user_input`. `.output` on the result is the `Contact`.
Raises:
UnexpectedModelBehavior: When the model can't produce a valid contact within
the retry budget. Callers decide what to do (retry later, queue for review).
"""
if deps is None:
deps = ExtractionDeps()
logger.info("Running extraction agent", extra={"user_input": user_input})
flow = Flow(USAGE_LIMITS)
result = await flow.run(extraction_agent, user_input, deps=deps)
return flow.finish(result.output)
if __name__ == "__main__":
import asyncio
configure_logging()
text = "Hi, it's Ada Lovelace from Analytical Engines Ltd. Reach me at ada@example.com."
print(asyncio.run(run_extraction(text)).output)
title = "Structured extraction"
pattern = "extraction"
summary = "Turn free text into a validated schema, with an output validator that sends bad answers back for correction."
smoke_input = "Hi, it's Ada Lovelace from Analytical Engines Ltd. Reach me at ada@example.com."
# TestModel's generated junk ("a" for every string) fails the validator, so the
# smoke tests supply an output it accepts.
[smoke.extraction_agent]
output = { name = "Ada Lovelace", email = "ada@example.com", company = "Analytical Engines Ltd" }
[entrypoint]
deps = "ExtractionDeps"
run = "run_extraction"
"""The output validator sends bad answers back to the model, within a retry budget."""
import pytest
from pydantic_ai.exceptions import UnexpectedModelBehavior
from pydantic_ai.messages import ModelResponse, RetryPromptPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
from examples.extraction.agent import ExtractionDeps, extraction_agent, run_extraction
def answers(*contacts: dict):
"""A FunctionModel that returns each contact in turn and records what it was sent."""
seen: list[list] = []
remaining = list(contacts)
def model_fn(messages, info: AgentInfo) -> ModelResponse:
seen.append(messages)
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, remaining.pop(0))])
return FunctionModel(model_fn), seen
async def test_a_valid_contact_is_returned_as_is():
model, seen = answers({"name": "Ada", "email": "ada@example.com"})
with extraction_agent.override(model=model):
result = await run_extraction("Ada, ada@example.com")
contact = result.output
assert contact.email == "ada@example.com"
assert len(seen) == 1
async def test_a_bad_email_is_sent_back_for_correction():
model, seen = answers(
{"name": "Ada", "email": "ada at example dot com"},
{"name": "Ada", "email": "ada@example.com"},
)
with extraction_agent.override(model=model):
result = await run_extraction("Ada, ada@example.com")
contact = result.output
assert contact.email == "ada@example.com"
assert len(seen) == 2
# One agent step, and the retry is visible in its usage: two model requests.
assert [step.agent for step in result.steps] == ["extraction"]
assert result.usage.requests == 2
retries = [p for p in seen[1][-1].parts if isinstance(p, RetryPromptPart)]
assert "is not an email address" in str(retries[0].content)
async def test_a_contact_with_no_way_to_reach_them_is_rejected():
model, _ = answers({"name": "Ada"}, {"name": "Ada", "phone": "+44 20 7946 0958"})
with extraction_agent.override(model=model):
result = await run_extraction("Ada, 020 7946 0958")
contact = result.output
assert contact.phone == "+44 20 7946 0958"
async def test_the_retry_budget_runs_out_instead_of_returning_something_invalid():
bad = {"name": "Ada", "email": "nope"}
model, seen = answers(bad, bad, bad)
with extraction_agent.override(model=model), pytest.raises(UnexpectedModelBehavior):
await extraction_agent.run("x", deps=ExtractionDeps())
assert len(seen) == 3 # the first attempt plus retries={"output": 2}
"""Live check: the model fills the schema, and bad input exhausts the retries. `pytest -m eval`."""
import pytest
from pydantic_ai.exceptions import UnexpectedModelBehavior
from evals.trace import traced_run
from examples.extraction import agent as module
from examples.live_support import assert_every_agent_ran, run_as_script
pytestmark = pytest.mark.eval
def has_digits(value: str | None) -> bool:
"""Whether a field holds a number. Models sometimes write the string "None" for an empty
field instead of null; that is a harmless quirk, while an invented number is the failure."""
return value is not None and any(ch.isdigit() for ch in value)
@pytest.fixture(scope="module")
async def with_email():
return await traced_run(
module.run_extraction,
"Hi, it's Ada Lovelace from Analytical Engines Ltd. Reach me at ada@example.com.",
)
@pytest.fixture(scope="module")
async def phone_only():
return await traced_run(module.run_extraction, "Reach Grace Hopper at +1 555 0100.")
async def test_a_contact_with_an_email_is_extracted_exactly(with_email):
contact = with_email.result.output
assert contact.name == "Ada Lovelace"
assert contact.email == "ada@example.com"
# The text reads "Analytical Engines Ltd. Reach me…", so a trailing period is faithful.
assert contact.company is not None and contact.company.startswith("Analytical Engines Ltd")
assert not has_digits(contact.phone) # not in the text, so not invented
async def test_a_phone_only_contact_is_accepted_without_an_email(phone_only):
contact = phone_only.result.output
assert contact.name == "Grace Hopper"
assert contact.email is None or "@" not in contact.email # no address in the text
assert contact.phone is not None and "555 0100" in contact.phone
async def test_text_with_no_way_to_reach_anyone_exhausts_the_retries_instead_of_inventing_one():
"""The validator rejects a contact with neither email nor phone; the model can't comply."""
with pytest.raises(UnexpectedModelBehavior, match="retries"):
await module.run_extraction("Lovely weather today, isn't it?")
async def test_each_run_is_one_agent_step(with_email, phone_only):
for traced in (with_email, phone_only):
assert [step.agent for step in traced.result.steps] == ["extraction"]
assert_every_agent_ran(module, with_email.agents_ran | phone_only.agents_ran)
async def test_the_demo_script_runs():
assert "ada@example.com" in await run_as_script("examples.extraction.agent")
Recorded run · gemini-3.1-flash-lite · 1 step · $0.0001
Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 1 step · 222 tokens · $0.0001 · 1.2 s.
Model output varies between runs. Regenerate with uv run python scripts/record_example.py extraction.
Input
Hi, it's Ada Lovelace from Analytical Engines Ltd. Reach me at ada@example.com.
Steps
1. extraction
222 tokens · $0.0001
Prompt
Hi, it's Ada Lovelace from Analytical Engines Ltd. Reach me at ada@example.com.
Output
{
"name": "Ada Lovelace",
"email": "ada@example.com",
"phone": null,
"company": "Analytical Engines Ltd."
}
Result
run_extraction(...).output