Skip to content

Evals

Each agent gets its own eval starter, evals/test_<name>.py, from add_agent.py:

  • A smoke eval, and a dataset eval driven by evals/fixtures/<name>.json (a pydantic_evals Dataset). Add cases to that JSON file to grow the eval; no code changes needed unless a case requires a new kind of check (then add an Evaluator alongside ContainsExpected in evals/helpers.py). Every case is also checked against behavioral budgets (MaxModelRequests, MaxToolCalls) read from the run's OpenTelemetry spans, so it grades how the agent got its answer, not just the answer. Tool-using agents can add optional keys to a fixture: expected_tools (["a", "b"], any order), expected_trajectory (ordered tool names, scored by F1) and expected_arguments ({"tool": "a", "args": {"q": "x"}}).
  • An LLM-as-judge eval, graded by AGENT_JUDGE_MODEL (see Configuration above); edit its criteria.

The shared evaluators and runner live in evals/helpers.py; the judge in evals/judge.py.

All of these share the same @pytest.mark.eval marker — there's no separate marker for the LLM-judge subset. uv run pytest -m eval runs all of them and requires a real API key; the LLM-judge evals also cost money (they make an extra model call per test to grade the output).