Evals¶
Each agent gets its own eval starter, evals/test_<name>.py, from add_agent.py:
- A smoke eval, and a dataset eval driven by
evals/fixtures/<name>.json(apydantic_evalsDataset). Add cases to that JSON file to grow the eval; no code changes needed unless a case requires a new kind of check (then add anEvaluatoralongsideContainsExpectedinevals/helpers.py). Every case is also checked against behavioral budgets (MaxModelRequests,MaxToolCalls) read from the run's OpenTelemetry spans, so it grades how the agent got its answer, not just the answer. Tool-using agents can add optional keys to a fixture:expected_tools(["a", "b"], any order),expected_trajectory(ordered tool names, scored by F1) andexpected_arguments({"tool": "a", "args": {"q": "x"}}). - An LLM-as-judge eval, graded by
AGENT_JUDGE_MODEL(see Configuration above); edit its criteria.
The shared evaluators and runner live in evals/helpers.py; the judge in evals/judge.py.
All of these share the same @pytest.mark.eval marker — there's no separate marker for the LLM-judge subset. uv run pytest -m eval runs all of them and requires a real API key; the LLM-judge evals also cost money (they make an extra model call per test to grade the output).