Guides
Tutorial: evaluate an agent
A runnable walkthrough for a tool-using agent and for a team of agents: record what the agent did as a trajectory, check its tool use, constraints, steps, routing, permissions and hand-offs, gate a release on them, and compare a candidate change with the baseline. Both examples run locally with no provider credentials.
The reference for every field and evaluator is Agents and tools; the terms case, evaluator, metric, interval and gate are in Core concepts. This page is the hands-on route through them.
What Oloproof does and does not do here
Oloproof does not drive your agent. Your application runs its own loop, calls its own tools, and records what happened as an agent_trajectory/v1 artifact. Every agent metric is read out of that record.
Your application also owns everything the tools touch. Oloproof provides no sandbox, no simulated tools and no reset between cases: if a tool writes to a database, sends an email or charges a card during an evaluation, it really does. Point the agent at test accounts, stubbed tools or a disposable environment, and reset state between cases yourself, before you run an evaluation.
Keep two kinds of question apart:
| Question | Checked by | Example |
|---|---|---|
| Did the user get the right outcome? (task success) | An output check such as contains, or a judge | answer_correct |
| Did the agent behave as allowed on the way? | Trajectory checks: tool choice, order, loops, constraints, steps, routing, permissions, hand-offs | agent_constraints_satisfied, agent_route |
They disagree in useful ways. In both examples below some cases answer correctly and still break a rule, and only a trajectory check sees it. A trajectory check passing says nothing about whether the task succeeded either.
Prerequisites
- Python 3.11 or later, and Oloproof installed (pip install oloproof).
- The example projects, which ship with the package: support_agent (one agent) and triage_agents (three). Copy one into a new directory and work there:
oloproof init --example support_agent my-agent
cd my-agentEvery command below runs from inside the copied directory. Evidence is stored in .oloproof/ there.
Part 1: one tool-using agent
The files
| File | What it is |
|---|---|
| app.py | The agent: Tools, a plan standing in for the model's decisions, and run(case), its loop, which records the trajectory |
| data/refunds.jsonl | 40 refund requests, each with the expected answer and, for most, the expected tools |
| data/orders.jsonl | The orders the lookup_order tool reads |
| oloproof.yaml | The suite: dataset, system, evaluators, distribution metrics, slices |
| release.yaml | The release policy |
Recording the trajectory
run is the whole integration surface. It calls each tool, appends an AgentStep for the call and one for its result, records the constraint checks its own environment made, and hands the trajectory to the case recorder:
@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
for name, arguments in plan(case):
steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
result = getattr(tools, name)(**arguments)
steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
...
current_case().agent_trajectory(
AgentTrajectory(
steps=tuple(steps),
terminal_status="success" if refunded else "failure",
truncated=truncated,
step_limit=STEP_LIMIT if truncated else None,
constraints=(AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),),
checkpoints=tuple(checkpoints),
)
)
return {"answer": "refunded" if refunded else "unresolved"}The artifact's shape:
| Field | What it records |
|---|---|
| steps | Each AgentStep: index, kind (message, tool_call, tool_result, observation, decision, final, or handoff), tool_name, arguments, result, and for teams agent and to_agent |
| terminal_status | success, failure or unknown, as the agent saw it |
| truncated, step_limit | That the loop hit its bound and the record stops short |
| constraints | AgentConstraintCheck(name, passed, step_index): checks your environment made, such as "no customer was deleted" |
| checkpoints | AgentCheckpoints a replay could resume from (see Limitations) |
To use your own agent, keep the recording and replace the loop: call your framework in run, and translate its steps into AgentStep as they happen. The system declares records: [agent_trajectory/v1] in oloproof.yaml; without it, the agent evaluators refuse to run rather than count every case as missing.
What a case declares
{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}expected.answer is for the task check. expected.tools is the tool sequence the case should follow; leave it out and tool-sequence checks do not apply to the case (it leaves their denominator rather than passing). behaviour is how this deterministic example chooses what its stand-in agent does; your cases carry only real inputs.
Choosing evaluators
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_tool_called, tool_name: lookup_order}
- {type: agent_no_tool_loop, max_repeats: 2}
- {type: agent_tool_sequence}
- {type: agent_constraints_satisfied, constraints: [no_deletion]}
- {type: agent_max_steps, max_steps: 10}
metrics:
- {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
- {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3- answer_correct is the task check.
- agent_tool_called asks for a required tool; agent_tool_sequence compares the calls with expected.tools; agent_no_tool_loop flags the same call repeated more than max_repeats times in a row. These describe tool use, not success.
- agent_constraints_satisfied reads the checks your environment recorded. Oloproof does not observe side effects itself, so a constraint your application does not record cannot be checked.
- agent_max_steps bounds each run; the two quantile metrics show the distribution, so a change that makes every run longer is visible before any one run hits the bound.
The release policy
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answer-floor
metric: answer_correct
min: 0.70
- id: tool-sequence-floor
metric: agent_tool_sequence
min: 0.70
- id: no-deletion
metric: agent_constraints_satisfied
kind: observed_count
max_failures: 0no-deletion is an observed-count rule: "this must not happen in the suite we ran" needs no interval. See Gating.
Run it
oloproof runRun run_01M4FCF6544JRDB16NJ1ZFPVRZ [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ tool-sequence-floor │ agent_tool_sequence │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-deletion │ agent_constraints_satisfied │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 82.5% │ [67.2%, 92.7%] │ 33 / 40 observed · 0 missing · 0 excluded │
│ agent_tool_lookup_order_called │ 100.0% │ [86.8%, 100.0%] │ 39 / 39 observed · 1 missing · 0 excluded │
│ agent_no_tool_loop │ 74.4% │ [56.1%, 87.4%] │ 29 / 39 observed · 1 missing · 0 excluded │
│ agent_tool_sequence │ 60.0% │ [43.3%, 75.2%] │ 24 / 40 observed · 0 missing · 0 excluded │
│ agent_constraints_satisfied │ 92.5% │ [79.6%, 98.5%] │ 37 / 40 observed · 0 missing · 0 excluded │
│ agent_steps_le_10 │ 97.5% │ [86.8%, 100.0%] │ 39 / 40 observed · 0 missing · 0 excluded │
│ steps_p95 │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50 │ 2 calls │ [2, 3] calls │ p50 of 39 observed · 1 missing · 0 excluded │
Cache: execution 0 hit/40 miss; judgment 0 hit/240 missHow to read it:
- Exit 1: a rule FAILed. Three cases called delete_customer, and the environment recorded the constraint as broken.
- answer-floor is INSUFFICIENT_EVIDENCE although 82.5% is above 70%: with 40 cases the interval still reaches 67.2%.
- 1 missing: one case hit the step bound, so its trace is truncated. A truncated trace proves some things (it did exceed 10 steps) and leaves others open (a required tool may be in the part not recorded), so those criteria count it as missing, and the interval allows it to have gone either way.
Inspect the failures
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case case_035The second prints one case in full. Trimmed:
case case_035
input: {
"order_id": "ord-036",
"behaviour": "violates"
}
output: {
"answer": "refunded"
}
judgments:
answer_correct: passed
agent_tool_lookup_order_called: passed
agent_no_tool_loop: passed
agent_tool_sequence: failed
agent_constraints_satisfied: failed
agent_steps_le_10: passedThe customer got the refund (task success) from an agent that deleted a customer on the way (a constraint broken). Neither result implies the other. The full trajectory, every step with its arguments and result, is in the exported bundle (oloproof export RUN_ID) and on the case view in the workbench. The meaningful next action is in the application: stop the loop from calling a tool it must never call.
Make a candidate change and compare
In app.py, make the loop refuse the forbidden tool:
for name, arguments in plan(case):
if name == FORBIDDEN:
continue # the candidate: the loop refuses the forbidden toolWrite a comparison policy, compare.yaml:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answers-not-worse
metric: answer_correct
kind: non_inferiority
margin: 0.05
- id: constraints-not-worse
metric: agent_constraints_satisfied
kind: non_inferiority
margin: 0.05oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlThe candidate run alone: no-deletion now PASSes, agent_constraints_satisfied reads 40 / 40, and the gate blocks with exit 3 because the two floors are still INSUFFICIENT_EVIDENCE. The comparison:
Comparison sha256:72836e90… of run_01M4FCFH0K2RRCN3ARAEN189X7 against run_01M4FCF6544JRDB16NJ1ZFPVRZ · 40 paired cases
answer_correct: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
agent_tool_lookup_order_called: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_no_tool_loop: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_tool_sequence: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_constraints_satisfied: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_steps_le_10: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
steps_p95: +0 steps [+0, no bound] steps · p95 of per-case differences · 39 paired · 1 missing
tool_calls_p50: +0 calls [+0, +0] calls · p50 of per-case differences · 39 paired · 1 missing
72 exploratory slice differences not shown; add --slices to list them
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
constraints-not-worse agent_constraints_satisfied non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 10 more paired cases would decide it, if the difference holds (50 in total at 8% discordance)
Gate: BLOCK (exit 3)Read the two halves separately. On this suite, the change removed every observed deletion, which the single-run observed-count rule already settles. Whether the candidate is not worse than the baseline in general is a different question, and 40 paired cases cannot yet establish it within a 5-point margin; the planning line says roughly how many more would. No answer changed, so task success is untouched by the fix.
Part 2: a team of agents
oloproof init --example triage_agents my-team
cd my-teamThe files and the recording
app.py runs three agents in one loop: triage hands each request to billing or tech, each specialist calls its own tools, and a refund billing may not issue is handed to a person. Each step names the agent that took it, and each transfer of control is a handoff step:
steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))A trajectory names the agent of every step or of none; one that names only some is refused. A case declares the route it should take:
{"id": "case_009", "input": {"topic": "tech", "request": "Two-factor codes are rejected", "order_id": "ord-009", "behaviour": "overreach"}, "expected": {"answer": "fixed", "route": ["triage", "tech"]}, "metadata": {"topic": "tech", "behaviour": "overreach"}}Evaluators and policy
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_route}
- type: agent_tool_permissions
permissions:
triage: []
billing: [lookup_order, issue_refund]
tech: [search_kb]
- {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
min_slice_support: 3- agent_route compares the agents that held control (repeats collapsed, a hand-off's receiver included) with expected.route. A routing check, not a success check.
- agent_tool_permissions checks every call against a closed map: an agent the map does not list may call no tool.
- agent_max_handoffs bounds how often control changed hands.
release.yaml has answer-floor (min: 0.80), routing-floor (min: 0.70) and no-overreach, an observed-count rule with max_failures: 0 on agent_tool_permissions.
Run the team
oloproof runGate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ PASS │ lower_bound_meets_minimum │
│ routing-floor │ agent_route │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-overreach │ agent_tool_permissions │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 96.7% │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1% │ [73.4%, 99.2%] │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2 │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │oloproof inspect RUN_ID --failures6 of 30 cases failed, errored or did not finish
case_005
output: {"answer": "refunded"}
agent_route: failed
agent_handoffs_le_2: failed
case_009
output: {"answer": "fixed"}
agent_tool_permissions: failed
...
case_030
output: {"answer": "unresolved"}
answer_correct: failed
agent_route: failed
agent_tool_permissions: error: MissingFieldError: truncated_trajectory: the trace stops before whether an agent called a tool it was not given is settled
agent_handoffs_le_2: failed- case_009 and case_020: tech issued a refund, a tool only billing holds. Both answered correctly. Task success, permission broken.
- case_005 and two others went to the wrong specialist first and came back through triage: the answer is right, the route and the hand-off bound are not.
- case_030 bounced between billing and tech until the loop's bound. Its truncated trace already proves the route and hand-off failures, and cannot settle permissions, so that criterion is missing for it rather than passed.
Nothing in the output says which agent is to blame. A route divergence says where two routes part; that an agent caused a failure is a claim about what would have happened had it acted otherwise, which no check here makes.
Change the team and compare
The meaningful next action for the permission failures: tech hands a refund to billing instead of issuing it. In app.py, in tech:
# The candidate: tech hands the refund to billing, the agent allowed to issue it.
trace.hand_off("tech", "billing", "a goodwill refund")
trace.call("billing", "issue_refund", order_id=str(case["order_id"]))With a compare.yaml holding answers-not-worse on answer_correct and routing-not-worse on agent_route, both non_inferiority with margin: 0.05:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlThe candidate run alone:
Gate: BLOCK (exit 3)
│ answer-floor │ answer_correct │ PASS │ lower_bound_meets_minimum │
│ routing-floor │ agent_route │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-overreach │ agent_tool_permissions │ INSUFFICIENT_EVIDENCE │ missing_could_change_outcome │
│ agent_route │ 80.0% │ [61.4%, 92.3%] │ 24 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 100.0% │ [82.7%, 100.0%] │ 29 / 29 observed · 1 missing · 0 excluded │and the comparison:
answer_correct: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
agent_route: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
agent_tool_permissions: +6.9 points [-18.9, +33.5] · 29 paired · 1 missing · 0 excluded
agent_handoffs_le_2: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
routing-not-worse agent_route non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
no sample size would make this PASS: the difference itself (-6.7 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)Three things to take from it:
- No observed call broke a permission, but no-overreach is now INSUFFICIENT_EVIDENCE rather than PASS: the truncated case_030 could hide a violation in the part not recorded (missing_could_change_outcome). Fixing that loop, not the permission map, is what would settle the rule.
- The fix changed the route of the two cases to triage > tech > billing, which their expected.route does not declare, so agent_route and agent_handoffs_le_2 fell. Whether that route is now correct is a product decision: if it is, update the cases' expected.route; a routing check measures conformance to what you declared, not quality.
- The planning line says more cases would move routing-not-worse toward FAIL, not PASS. The comparison is telling you the candidate as written trades routing for permissions.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Agent evaluators refuse to run | records: [agent_trajectory/v1] missing on the system | Declare it in oloproof.yaml and on @system |
| A trajectory is refused | Some steps name an agent and others do not | Name the agent of every step, or of none |
| Many cases missing on a criterion | Truncated traces: the loop hit its bound | Raise the bound, or fix the loop; missing cases widen the interval rather than pass |
| agent_tool_sequence has a small denominator | Cases without expected.tools | Declare the sequence where it matters; [] means "expects no tool" |
| A constraint metric never fails | The application does not record that check | Record an AgentConstraintCheck where your environment observes it |
| Results differ between runs of the same version | Tools read or write shared state | Reset that state before each case in your application; Oloproof does not |
| An agent added to the team fails permissions at once | The permission map is closed | Declare what the new agent may call |
Limitations
- Oloproof does not drive, sandbox or reset an agent. Tool side effects, sessions, state and their reset belong to your application.
- Every check reads the recorded trajectory. What the application does not record cannot be measured, and a truncated trace counts as missing wherever its prefix does not settle the question.
- Trajectory checks are deterministic rules. There is no LLM-judged trajectory quality check.
- No output attributes a failure to a step or an agent. Agent replay, which re-runs a case from a recorded checkpoint with a step dropped to label it needed or unnecessary, exists in the Python SDK only (replay_case), for a system that implements replay from its checkpoints; there is no CLI command for it, and nothing above uses it.
- Multi-turn conversations are a different surface (SDK only); see What works today.
- The examples' plan and behaviour fields stand in for a model's decisions so the runs are reproducible. A live model in your loop calls a provider, needs credentials and costs money per case.