Human in the loop¶
Pause a risky action until a person approves it, reject impossible requests before anyone is asked, and resume the same run with the decision.
Some actions an agent can take should not happen on the model's say-so alone: a refund, a deletion, a message sent in your name. Here a tool can pause the run. Instead of acting, it asks for approval and the run stops. A person (or a policy, or a queue you build) decides, and the same run resumes with that decision; if it was denied, the model is told, so it can explain. Requests that can't succeed are rejected first, so nobody is asked about something impossible, and the agent's own claims are checked against a ledger of what really happened.
Use it when
- The agent can take actions with real consequences: money, deletions, messages.
- Some of them should wait for a human and others are fine on their own.
- A request that can't succeed shouldn't cost anyone's attention.
Look elsewhere when
- Every action is safe or reversible: the pause only adds delay.
- You need to filter what users send in, not approve what the agent does:
guardrails.
request → tool call → validate
├─ impossible → rejected; no human is asked
├─ small → runs
└─ large → PAUSE → approver → resume → answer
What it shows
- Three outcomes from one tool (
issue_refund): refunds up toAUTO_APPROVE_LIMIT_USDrun on their own; larger ones raiseApprovalRequiredand the run pauses, returningDeferredToolRequestsinstead of acting; impossible ones (unknown order, shipped, over the total) are rejected first - Validation before approval:
args_validatorruns before the approval gate, so a human is only ever asked about a request that could really happen - A pluggable approver:
run_refundsasksdeps.approverabout each pending call, then resumes the same conversation withDeferredToolResults. A denial is passed to the model as aToolDeniedmessage so it can explain.deny_allis the default (nothing risky happens unconfigured);console_approverasks a person at the terminal; swap in a Slack message, a ticket queue or a policy - The ledger is the truth:
deps.refundsrecords what really happened, and an output validator rejects a claim that doesn't match it (refunded=Truewith an empty ledger is sent back to the model), so the agent can't say it did something it didn't - A run that keeps asking for approval is stopped after
MAX_APPROVAL_ROUNDS
See it run: sample_run.md is a recorded run against a real model: what each agent was asked, which tools it called, and what it returned.
To adapt it, replace the orders and issue_refund with your own action, keep the validate → defer →
resume shape, and replace approver with how your approvals really happen. The result has one step
for the run that paused and one for each run that resumed.
Source¶
All of it is in examples/human_in_the_loop/.
"""Human in the loop: pause a risky action until a person approves it.
Use this pattern when:
- The agent can take actions with real consequences (money, deletions, messages)
- Some of those actions should wait for a human, and others are fine to run alone
- Impossible requests should be rejected without bothering anyone
How it works:
1. `issue_refund` is a tool whose arguments are checked first (`args_validator`): an unknown
order, a shipped order or an amount over the total is rejected on the spot, and the human is
never asked about it
2. A refund above AUTO_APPROVE_LIMIT_USD raises `ApprovalRequired`: the run *pauses* and returns
`DeferredToolRequests` instead of acting
3. `run_refunds` asks an approver (a human, a policy, a ticket queue) about each request, then
resumes the same run with their decisions; a denial is explained to the model
4. An output validator checks that what the model says happened is what the ledger shows
`deps.refunds` is the ledger and the ground truth for whether anything really happened, which is
what the tests check, rather than trusting what the model says.
"""
from __future__ import annotations
import asyncio
from collections.abc import Awaitable, Callable
from dataclasses import dataclass, field
from pydantic import BaseModel
from pydantic_ai import (
Agent,
ApprovalRequired,
DeferredToolRequests,
DeferredToolResults,
ModelRetry,
RunContext,
ToolDenied,
ToolFailed,
)
from pydantic_ai.capabilities import RaiseContentFilterError
from pydantic_ai.messages import ToolCallPart
from pydantic_ai.usage import UsageLimits
from agent.config import settings
from agent.logging import agent_label, configure_logging, get_logger
from agent.prompts.templates import load_prompt
from agent.runs import Flow, RunResult
logger = get_logger(__name__)
LABEL = agent_label(__name__) # names this agent's run spans in Logfire traces
# Each approval round is another model request, so leave headroom for a couple of rounds.
USAGE_LIMITS = UsageLimits(
request_limit=12, total_tokens_limit=100_000, cost_limit=settings.cost_limit
)
AUTO_APPROVE_LIMIT_USD = 25.0 # refunds up to this run without a human
MAX_APPROVAL_ROUNDS = 3 # a run that keeps asking for approval is stopped
# --- The business data ---
@dataclass(frozen=True)
class Order:
id: str
item: str
total: float
status: str # "delivered" or "shipped"
ORDERS: dict[str, Order] = {
"A100": Order("A100", "Trailhead boots", 84.50, "delivered"),
"A101": Order("A101", "Wool socks", 18.00, "delivered"),
"A102": Order("A102", "Summit pack", 129.00, "shipped"),
}
@dataclass(frozen=True)
class Refund:
order_id: str
amount: float
reason: str
# --- Approvers: how a decision gets made ---
# An approver is shown the pending tool call and returns True to approve, or a string explaining
# why it is denied (the model is told that string).
Approver = Callable[[ToolCallPart], Awaitable[bool | str]]
async def deny_all(call: ToolCallPart) -> bool | str:
"""The safe default: with no approver configured, nothing risky happens."""
return "No approver is configured, so this action was not approved."
async def approve_all(call: ToolCallPart) -> bool | str:
"""Approve everything. For demos and tests; never as a production default."""
return True
async def console_approver(call: ToolCallPart) -> bool | str:
"""Ask a person at the terminal."""
args = call.args_as_dict()
print(f"\nApproval needed: {call.tool_name}({args})")
answer = await asyncio.to_thread(input, "Approve? [y/N] ")
return True if answer.strip().lower() in {"y", "yes"} else "The reviewer declined."
# --- Dependencies ---
@dataclass
class RefundDeps:
"""Runtime dependencies for the refunds agent."""
orders: dict[str, Order] = field(default_factory=lambda: dict(ORDERS))
approver: Approver = deny_all
refunds: list[Refund] = field(default_factory=list) # the ledger: what really happened
decisions: list[tuple[str, bool]] = field(default_factory=list) # (order id, approved?) asked
# --- Output type ---
class RefundOutput(BaseModel):
# `result` is the conventional output field in these examples; the generated
# eval starter reads it when present (see evals/helpers.py).
result: str
refunded: bool
# --- Agent ---
# DeferredToolRequests in the output type is what lets a run pause for approval.
refund_agent: Agent[RefundDeps, RefundOutput | DeferredToolRequests] = Agent(
settings.model,
name=LABEL,
output_type=[RefundOutput, DeferredToolRequests],
deps_type=RefundDeps,
capabilities=[RaiseContentFilterError()],
instructions=load_prompt("human_in_the_loop"), # prompts/…; copied to agent/prompts/<name>.txt
)
# --- Tools ---
@refund_agent.tool
async def lookup_order(ctx: RunContext[RefundDeps], order_id: str) -> str:
"""Look up an order: its item, total and status.
Args:
order_id: The order id, such as "A100".
Raises:
ToolFailed: When there is no such order.
"""
order = ctx.deps.orders.get(order_id.strip().upper())
if order is None:
raise ToolFailed(f"There is no order {order_id!r}. Ask the customer to check the number.")
return f"Order {order.id}: {order.item}, total ${order.total:.2f}, status {order.status}."
def check_refund(
ctx: RunContext[RefundDeps], order_id: str, amount_usd: float, reason: str
) -> None:
"""Runs before the tool and before any approval, so a human is only asked about real requests.
Raises ToolFailed (nothing to find or do), ModelRetry (the model can correct the amount), or
ApprovalRequired (valid, but a person must say yes).
"""
order = ctx.deps.orders.get(order_id.strip().upper())
if order is None:
raise ToolFailed(f"There is no order {order_id!r}.")
if order.status != "delivered":
raise ToolFailed(
f"Order {order.id} is {order.status}, not delivered, so it cannot be refunded."
)
if any(refund.order_id == order.id for refund in ctx.deps.refunds):
raise ToolFailed(f"Order {order.id} has already been refunded.")
if not 0 < amount_usd <= order.total:
raise ModelRetry(
f"The refund must be more than $0 and at most the order total of ${order.total:.2f}."
)
if amount_usd > AUTO_APPROVE_LIMIT_USD and not ctx.tool_call_approved:
raise ApprovalRequired() # pauses the run; the tool does not execute yet
@refund_agent.tool(args_validator=check_refund)
async def issue_refund(
ctx: RunContext[RefundDeps], order_id: str, amount_usd: float, reason: str
) -> str:
"""Refund part or all of a delivered order. Larger refunds wait for a person's approval.
Args:
order_id: The order id, such as "A100".
amount_usd: The amount to refund, at most the order total.
reason: Why the customer is asking, in a few words.
"""
refund = Refund(order_id.strip().upper(), amount_usd, reason)
ctx.deps.refunds.append(refund) # the only place money moves
logger.info("Refund issued", extra={"order": refund.order_id, "amount": amount_usd})
return f"Refunded ${amount_usd:.2f} on order {refund.order_id}."
# --- Honest reporting ---
@refund_agent.output_validator
def check_the_claim_matches_the_ledger(
ctx: RunContext[RefundDeps], output: RefundOutput | DeferredToolRequests
) -> RefundOutput | DeferredToolRequests:
"""The model may not say a refund happened unless the ledger shows one (or vice versa)."""
if isinstance(output, RefundOutput) and output.refunded != bool(ctx.deps.refunds):
raise ModelRetry(
f"You set refunded={output.refunded}, but issue_refund has "
f"{'run' if ctx.deps.refunds else 'not run'}. Report what actually happened."
)
return output
class ApprovalLoopError(Exception):
"""The agent kept asking for approval instead of finishing."""
async def run_refunds(user_input: str, deps: RefundDeps | None = None) -> RunResult[RefundOutput]:
"""Handle a refund request, pausing for approval where one is needed.
Returns:
A RunResult: `.output` is the final `RefundOutput`. `.steps` has one step per run of the
agent: the first, then one for each time it resumed after an approval round.
Raises:
ApprovalLoopError: When the agent is still waiting for approval after
MAX_APPROVAL_ROUNDS rounds.
"""
if deps is None:
deps = RefundDeps()
logger.info("Running refunds agent", extra={"user_input": user_input})
flow = Flow(USAGE_LIMITS)
result = await flow.run(refund_agent, user_input, deps=deps)
for _ in range(MAX_APPROVAL_ROUNDS):
if not isinstance(result.output, DeferredToolRequests):
return flow.finish(result.output)
decisions = DeferredToolResults()
for call in result.output.approvals:
verdict = await deps.approver(call)
approved = verdict is True
deps.decisions.append((str(call.args_as_dict().get("order_id")), approved))
decisions.approvals[call.tool_call_id] = True if approved else ToolDenied(str(verdict))
# Resume the same conversation with the decisions; there is no new prompt.
result = await flow.run(
refund_agent,
None,
deps=deps,
message_history=result.all_messages(),
deferred_tool_results=decisions,
)
if isinstance(result.output, DeferredToolRequests):
raise ApprovalLoopError(f"Still waiting for approval after {MAX_APPROVAL_ROUNDS} rounds")
return flow.finish(result.output)
if __name__ == "__main__":
import sys
configure_logging()
# A person at a terminal is asked; piped or scripted, approve so the demo completes.
approver = console_approver if sys.stdin.isatty() else approve_all
demo = RefundDeps(approver=approver)
result = asyncio.run(
run_refunds("Order A100 arrived with a broken sole. Refund the full $84.50.", demo)
)
print(result.output, "| ledger:", demo.refunds)
You are a refunds assistant for Birchwood Outfitters.
- Look up the order with lookup_order, then call issue_refund for the amount the customer asks
for. If they ask for a full refund, use the order total.
- Some refunds need a human's approval before they happen. If a refund is declined or cannot be
processed, tell the customer plainly why. Do not promise it will happen anyway, and do not say you
have submitted, escalated or forwarded anything: you cannot do any of those.
- Set `refunded` to true only if issue_refund succeeded. Never claim a refund happened otherwise.
- If the customer only asks about an order, answer from lookup_order and do not issue a refund.
title = "Human in the loop"
pattern = "human_in_the_loop"
summary = "Pause a risky tool call until a person approves it; reject impossible requests before anyone is asked; resume the same run."
smoke_input = "Order A100 arrived with a broken sole. Please refund the full $84.50."
# A real model must look the order up for this (checked by the release check).
expected_tools = ["lookup_order"]
# TestModel's generated output claims refunded=true, which the honesty validator rejects while the
# ledger is empty, so the smoke tests supply an output that matches it.
[smoke.refund_agent]
output = { result = "I could not process that here.", refunded = false }
[entrypoint]
deps = "RefundDeps"
run = "run_refunds"
"""The refund tool's validation, the approval pause/resume loop, and the honesty check — offline.
The ledger (`deps.refunds`) is the ground truth throughout: these tests check what really happened,
not what the model said.
"""
import builtins
import pytest
from pydantic_ai import ApprovalRequired, DeferredToolRequests, ModelRetry, RunContext, ToolFailed
from pydantic_ai.messages import (
ModelResponse,
RetryPromptPart,
ToolCallPart,
ToolReturnPart,
)
from pydantic_ai.models.function import AgentInfo, FunctionModel
from pydantic_ai.models.test import TestModel
from pydantic_ai.usage import RunUsage
from examples.human_in_the_loop import agent as module
from examples.human_in_the_loop.agent import (
AUTO_APPROVE_LIMIT_USD,
ORDERS,
ApprovalLoopError,
Refund,
RefundDeps,
RefundOutput,
approve_all,
check_refund,
check_the_claim_matches_the_ledger,
console_approver,
deny_all,
issue_refund,
lookup_order,
refund_agent,
run_refunds,
)
def ctx(deps: RefundDeps | None = None, approved: bool = False) -> RunContext[RefundDeps]:
return RunContext(
deps=deps or RefundDeps(), model=TestModel(), usage=RunUsage(), tool_call_approved=approved
)
# --- lookup_order ---
async def test_an_order_is_described_whatever_the_case_of_its_id():
text = await lookup_order(ctx(), " a100 ")
assert text == "Order A100: Trailhead boots, total $84.50, status delivered."
async def test_an_unknown_order_is_a_terminal_failure():
with pytest.raises(ToolFailed, match="no order 'Z9'"):
await lookup_order(ctx(), "Z9")
# --- check_refund: runs before the tool and before any approval ---
def check(order="A100", amount=84.50, deps=None, approved=False):
check_refund(ctx(deps, approved), order, amount, "broken")
def test_an_unknown_order_is_rejected_before_anyone_is_asked():
with pytest.raises(ToolFailed, match="no order"):
check(order="nope")
def test_an_order_that_has_not_been_delivered_cannot_be_refunded():
with pytest.raises(ToolFailed, match="shipped, not delivered"):
check(order="A102", amount=10)
def test_an_order_cannot_be_refunded_twice():
deps = RefundDeps(refunds=[Refund("A101", 18.0, "x")])
with pytest.raises(ToolFailed, match="already been refunded"):
check(order="A101", amount=5, deps=deps)
@pytest.mark.parametrize("amount", [0, -5, 84.51, 500])
def test_an_amount_outside_zero_to_the_total_asks_the_model_to_correct_it(amount):
with pytest.raises(ModelRetry, match=r"at most the order total of \$84.50"):
check(amount=amount)
def test_a_small_refund_needs_no_approval():
check(order="A100", amount=AUTO_APPROVE_LIMIT_USD) # at the limit is still automatic
def test_a_larger_refund_pauses_for_approval():
with pytest.raises(ApprovalRequired):
check(amount=AUTO_APPROVE_LIMIT_USD + 0.01)
def test_once_approved_the_same_refund_passes():
check(amount=84.50, approved=True)
# --- issue_refund ---
async def test_issuing_a_refund_records_it_in_the_ledger():
deps = RefundDeps()
text = await issue_refund(ctx(deps), " a101", 18.0, "hole")
assert text == "Refunded $18.00 on order A101."
assert deps.refunds == [Refund("A101", 18.0, "hole")]
def test_a_deps_object_gets_its_own_copy_of_the_orders_and_denies_by_default():
deps = RefundDeps()
deps.orders.pop("A100")
assert "A100" in ORDERS # the shared table is untouched
assert deps.approver is deny_all
# --- Approvers ---
async def test_the_default_approver_declines_with_a_reason_the_model_can_relay():
verdict = await deny_all(ToolCallPart("issue_refund", {}))
assert isinstance(verdict, str) and "not approved" in verdict
async def test_approve_all_approves():
assert await approve_all(ToolCallPart("issue_refund", {})) is True
@pytest.mark.parametrize(
("typed", "approved"), [("y", True), (" YES ", True), ("n", False), ("", False)]
)
async def test_the_console_approver_asks_a_person(monkeypatch, capsys, typed, approved):
monkeypatch.setattr(builtins, "input", lambda prompt: typed)
call = ToolCallPart("issue_refund", {"order_id": "A100", "amount_usd": 84.5})
verdict = await console_approver(call)
assert (verdict is True) == approved
if not approved:
assert verdict == "The reviewer declined."
assert "Approval needed: issue_refund" in capsys.readouterr().out
# --- Honest reporting ---
def claim(refunded: bool, ledger: bool):
deps = RefundDeps(refunds=[Refund("A100", 1.0, "x")] if ledger else [])
return check_the_claim_matches_the_ledger(
ctx(deps), RefundOutput(result="r", refunded=refunded)
)
def test_a_claim_that_matches_the_ledger_passes():
assert claim(True, True).refunded and not claim(False, False).refunded
def test_claiming_a_refund_that_never_happened_is_sent_back():
with pytest.raises(ModelRetry, match="refunded=True, but issue_refund has not run"):
claim(True, False)
def test_denying_a_refund_that_did_happen_is_sent_back():
with pytest.raises(ModelRetry, match="refunded=False, but issue_refund has run"):
claim(False, True)
def test_a_pending_approval_request_passes_straight_through():
pending = DeferredToolRequests()
assert check_the_claim_matches_the_ledger(ctx(), pending) is pending
# --- The pause/resume loop, with a scripted model ---
def scripted(*steps: dict, seen: list | None = None):
"""Follow `steps` (a tool call, or the final output); optionally record what each call saw."""
remaining = list(steps)
def model_fn(messages, info: AgentInfo) -> ModelResponse:
if seen is not None:
seen.append(messages)
step = remaining.pop(0)
if "tool" in step:
return ModelResponse(parts=[ToolCallPart(step["tool"], step["args"])])
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, step["output"])])
return FunctionModel(model_fn)
def refund_call(order="A100", amount=84.50):
return {
"tool": "issue_refund",
"args": {"order_id": order, "amount_usd": amount, "reason": "damaged"},
}
def done(refunded: bool, text="Done."):
return {"output": {"result": text, "refunded": refunded}}
async def test_an_approved_refund_pauses_then_happens():
deps = RefundDeps(approver=approve_all)
model = scripted(refund_call(), done(True, "Refunded."))
with refund_agent.override(model=model):
result = await run_refunds("Refund A100", deps)
assert deps.refunds == [Refund("A100", 84.50, "damaged")] # it really happened
assert deps.decisions == [("A100", True)]
assert result.output == RefundOutput(result="Refunded.", refunded=True)
assert len(result.steps) == 2 # the run that paused, and the one that resumed
async def test_a_denied_refund_does_not_happen_and_the_model_is_told_why():
async def decline(call):
return "Refunds over $25 need a manager."
deps = RefundDeps(approver=decline)
seen: list = []
model = scripted(refund_call(), done(False, "A manager must approve that."), seen=seen)
with refund_agent.override(model=model):
result = await run_refunds("Refund A100", deps)
assert deps.refunds == [] # nothing happened
assert deps.decisions == [("A100", False)]
assert result.output.refunded is False
denied = [
p
for m in seen[-1]
for p in m.parts
if isinstance(p, ToolReturnPart) and p.outcome == "denied"
]
assert denied and "need a manager" in str(denied[0].content) # the resumed model saw the reason
async def test_a_small_refund_runs_straight_through_without_asking_anyone():
deps = RefundDeps(approver=pytest.fail) # would blow up if anyone were asked
model = scripted(refund_call("A101", 18.0), done(True, "Refunded $18."))
with refund_agent.override(model=model):
result = await run_refunds("Refund the socks", deps)
assert deps.refunds == [Refund("A101", 18.0, "damaged")]
assert deps.decisions == []
assert len(result.steps) == 1
async def test_an_impossible_refund_is_corrected_before_any_human_is_asked():
deps = RefundDeps(approver=pytest.fail)
model = scripted(refund_call(amount=500), done(False, "That is more than the order total."))
with refund_agent.override(model=model):
result = await run_refunds("Refund $500", deps)
retries = [p for m in result.all_messages() for p in m.parts if isinstance(p, RetryPromptPart)]
assert len(retries) == 1 and "order total" in str(retries[0].content)
assert deps.refunds == [] and deps.decisions == []
async def test_a_read_only_request_never_reaches_the_approval_gate():
deps = RefundDeps(approver=pytest.fail)
model = scripted(
{"tool": "lookup_order", "args": {"order_id": "A102"}}, done(False, "It has shipped.")
)
with refund_agent.override(model=model):
result = await run_refunds("Where is A102?", deps)
assert deps.refunds == [] and len(result.steps) == 1
async def test_a_model_that_claims_a_refund_that_did_not_happen_is_made_to_correct_itself():
model = scripted(done(True, "All refunded!"), done(False, "Actually nothing was refunded."))
with refund_agent.override(model=model):
result = await run_refunds("Refund A100", RefundDeps())
assert result.output.refunded is False # the first, false claim never got out
async def test_a_model_that_keeps_asking_is_stopped_after_the_round_limit():
asks = [refund_call() for _ in range(module.MAX_APPROVAL_ROUNDS + 1)]
with refund_agent.override(model=scripted(*asks)), pytest.raises(ApprovalLoopError):
await run_refunds("Refund A100", RefundDeps()) # default approver declines every time
async def test_the_last_permitted_round_can_still_finish():
asks = [refund_call() for _ in range(module.MAX_APPROVAL_ROUNDS)]
with refund_agent.override(model=scripted(*asks, done(False, "Declined."))):
result = await run_refunds("Refund A100", RefundDeps())
assert result.output.refunded is False and len(result.steps) == module.MAX_APPROVAL_ROUNDS + 1
"""Live check: approvals pause and resume a real run, and the ledger shows what really happened.
The ledger (`deps.refunds`) is the ground truth. What the model *says* is checked against it, never
the other way round. Run with `pytest -m eval`.
"""
import pytest
from evals.trace import traced_run
from examples.human_in_the_loop import agent as module
from examples.live_support import assert_every_agent_ran, run_as_script
pytestmark = pytest.mark.eval
async def decline(call):
return "Refunds over $25 need a manager."
async def run(prompt: str, approver):
"""Run the agent with a fresh ledger; return the traced run and the deps it used."""
deps = module.RefundDeps(approver=approver)
async def helper(text: str):
return await module.run_refunds(text, deps)
return await traced_run(helper, prompt), deps
@pytest.fixture(scope="module")
async def approved():
return await run(
"Order A100 arrived with a broken sole. Please refund the full $84.50.", module.approve_all
)
@pytest.fixture(scope="module")
async def denied():
return await run(
"Order A100 arrived with a broken sole. Please refund the full $84.50.", decline
)
@pytest.fixture(scope="module")
async def small():
return await run(
"Order A101 had a hole in the socks. Please refund the $18.00.", module.deny_all
)
@pytest.fixture(scope="module")
async def impossible():
return await run("Please refund $500 on order A100.", module.approve_all)
@pytest.fixture(scope="module")
async def read_only():
return await run("What is the status of order A102?", module.approve_all)
async def test_an_approved_refund_really_happens_after_the_run_pauses_and_resumes(approved):
traced, deps = approved
assert deps.refunds and deps.refunds[0].order_id == "A100"
assert deps.refunds[0].amount == pytest.approx(84.50)
assert deps.decisions == [("A100", True)] # a person was asked exactly once, and said yes
assert len(traced.result.steps) == 2 # the run that paused, and the one that resumed
assert traced.result.output.refunded is True and traced.result.output.result.strip()
async def test_a_denied_refund_does_not_happen_and_the_customer_is_told(denied):
traced, deps = denied
assert deps.refunds == [] # money did not move
assert deps.decisions == [("A100", False)]
assert traced.result.output.refunded is False # and the model reports that honestly
answer = traced.result.output.result.lower()
assert answer.strip()
# It must not claim an action it has no way to take (the prompt forbids it).
assert not [w for w in ("submitted", "escalated", "forwarded") if w in answer], answer
async def test_a_small_refund_goes_through_without_asking_a_person(small):
traced, deps = small
assert [r.order_id for r in deps.refunds] == ["A101"]
assert deps.refunds[0].amount == pytest.approx(18.00)
assert deps.decisions == [] # the always-deny approver was never consulted
assert len(traced.result.steps) == 1 and traced.result.output.refunded is True
async def test_an_impossible_refund_is_rejected_before_a_person_is_asked(impossible):
traced, deps = impossible
assert deps.refunds == [] and deps.decisions == [] # nobody was bothered, nothing moved
assert traced.result.output.refunded is False
assert len(traced.result.steps) == 1
async def test_a_question_about_an_order_changes_nothing(read_only):
traced, deps = read_only
assert deps.refunds == [] and deps.decisions == []
assert "lookup_order" in traced.tools_called
assert "shipped" in traced.result.output.result.lower()
assert traced.result.output.refunded is False
async def test_no_run_ever_claimed_something_the_ledger_contradicts(
approved, denied, small, impossible, read_only
):
"""The honesty check is enforced in code; this confirms it held across every real run."""
for traced, deps in (approved, denied, small, impossible, read_only):
assert traced.result.output.refunded == bool(deps.refunds)
async def test_the_agent_ran(approved, denied, small, impossible, read_only):
ran = set().union(
*(traced.agents_ran for traced, _ in (approved, denied, small, impossible, read_only))
)
assert_every_agent_ran(module, ran)
async def test_the_demo_script_runs():
out = await run_as_script("examples.human_in_the_loop.agent")
assert "ledger: [Refund(order_id='A100'" in out # stdin isn't a terminal, so it auto-approves
Recorded run · gemini-3.1-flash-lite · 2 steps · $0.0009
Recorded 2026-10-07 with google:gemini-3.1-flash-lite · 2 steps · 1,624 tokens · $0.0009 · 3.2 s.
Model output varies between runs. Regenerate with uv run python scripts/record_example.py human_in_the_loop.
Input
Order A100 arrived with a broken sole. Please refund the full $84.50.
Steps
1. human_in_the_loop
1,009 tokens · $0.0003
Prompt
Order A100 arrived with a broken sole. Please refund the full $84.50.
What happened
- called lookup_order({"order_id": "A100"})
- lookup_order returned: Order A100: Trailhead boots, total $84.50, status delivered.
- called issue_refund({"reason": "broken sole", "order_id": "A100", "amount_usd": 84.5})
Output
The run paused: waiting for a decision on
approve issue_refund({"reason": "broken sole", "order_id": "A100", "amount_usd": 84.5})
2. human_in_the_loop
1,624 tokens · $0.0005
Prompt
Order A100 arrived with a broken sole. Please refund the full $84.50.
What happened
- called lookup_order({"order_id": "A100"})
- lookup_order returned: Order A100: Trailhead boots, total $84.50, status delivered.
- called issue_refund({"reason": "broken sole", "order_id": "A100", "amount_usd": 84.5})
- issue_refund returned: No approver is configured, so this action was not approved.
Output
{
"result": "I'm sorry, but I was unable to process the refund for order A100 because it could not be approved.",
"refunded": false
}
Result
run_refunds(...).output