Skip to content

Guides

Tutorial: evaluate an agent

A runnable walkthrough for a tool-using agent and for a team of agents: record what the agent did as a trajectory, check its tool use, constraints, steps, routing, permissions and hand-offs, gate a release on them, and compare a candidate change with the baseline. Both examples run locally with no provider credentials.

The reference for every field and evaluator is Agents and tools; the terms case, evaluator, metric, interval and gate are in Core concepts. This page is the hands-on route through them.

What Oloproof does and does not do here

Oloproof does not drive your agent. Your application runs its own loop, calls its own tools, and records what happened as an agent_trajectory/v1 artifact. Every agent metric is read out of that record.

Your application also owns everything the tools touch. Oloproof provides no sandbox, no simulated tools and no reset between cases: if a tool writes to a database, sends an email or charges a card during an evaluation, it really does. Point the agent at test accounts, stubbed tools or a disposable environment, and reset state between cases yourself, before you run an evaluation.

Keep two kinds of question apart:

QuestionChecked byExample
Did the user get the right outcome? (task success)An output check such as contains, or a judgeanswer_correct
Did the agent behave as allowed on the way?Trajectory checks: tool choice, order, loops, constraints, steps, routing, permissions, hand-offsagent_constraints_satisfied, agent_route

They disagree in useful ways. In both examples below some cases answer correctly and still break a rule, and only a trajectory check sees it. A trajectory check passing says nothing about whether the task succeeded either.

Prerequisites

  • Python 3.11 or later, and Oloproof installed (pip install oloproof).
  • The example projects, which ship with the package: support_agent (one agent) and triage_agents (three). Copy one into a new directory and work there:
oloproof init --example support_agent my-agent
cd my-agent

Every command below runs from inside the copied directory. Evidence is stored in .oloproof/ there.

Part 1: one tool-using agent

The files

FileWhat it is
app.pyThe agent: Tools, a plan standing in for the model's decisions, and run(case), its loop, which records the trajectory
data/refunds.jsonl40 refund requests, each with the expected answer and, for most, the expected tools
data/orders.jsonlThe orders the lookup_order tool reads
oloproof.yamlThe suite: dataset, system, evaluators, distribution metrics, slices
release.yamlThe release policy

Recording the trajectory

run is the whole integration surface. It calls each tool, appends an AgentStep for the call and one for its result, records the constraint checks its own environment made, and hands the trajectory to the case recorder:

@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
    ...
    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),),
            checkpoints=tuple(checkpoints),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

The artifact's shape:

FieldWhat it records
stepsEach AgentStep: index, kind (message, tool_call, tool_result, observation, decision, final, or handoff), tool_name, arguments, result, and for teams agent and to_agent
terminal_statussuccess, failure or unknown, as the agent saw it
truncated, step_limitThat the loop hit its bound and the record stops short
constraintsAgentConstraintCheck(name, passed, step_index): checks your environment made, such as "no customer was deleted"
checkpointsAgentCheckpoints a replay could resume from (see Limitations)

To use your own agent, keep the recording and replace the loop: call your framework in run, and translate its steps into AgentStep as they happen. The system declares records: [agent_trajectory/v1] in oloproof.yaml; without it, the agent evaluators refuse to run rather than count every case as missing.

What a case declares

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.answer is for the task check. expected.tools is the tool sequence the case should follow; leave it out and tool-sequence checks do not apply to the case (it leaves their denominator rather than passing). behaviour is how this deterministic example chooses what its stand-in agent does; your cases carry only real inputs.

Choosing evaluators

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
  • answer_correct is the task check.
  • agent_tool_called asks for a required tool; agent_tool_sequence compares the calls with expected.tools; agent_no_tool_loop flags the same call repeated more than max_repeats times in a row. These describe tool use, not success.
  • agent_constraints_satisfied reads the checks your environment recorded. Oloproof does not observe side effects itself, so a constraint your application does not record cannot be checked.
  • agent_max_steps bounds each run; the two quantile metrics show the distribution, so a change that makes every run longer is visible before any one run hits the bound.

The release policy

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: tool-sequence-floor
    metric: agent_tool_sequence
    min: 0.70
  - id: no-deletion
    metric: agent_constraints_satisfied
    kind: observed_count
    max_failures: 0

no-deletion is an observed-count rule: "this must not happen in the suite we ran" needs no interval. See Gating.

Run it

oloproof run
Run run_01M4FCF6544JRDB16NJ1ZFPVRZ [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor        │ answer_correct              │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ tool-sequence-floor │ agent_tool_sequence         │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
Cache: execution 0 hit/40 miss; judgment 0 hit/240 miss

How to read it:

  • Exit 1: a rule FAILed. Three cases called delete_customer, and the environment recorded the constraint as broken.
  • answer-floor is INSUFFICIENT_EVIDENCE although 82.5% is above 70%: with 40 cases the interval still reaches 67.2%.
  • 1 missing: one case hit the step bound, so its trace is truncated. A truncated trace proves some things (it did exceed 10 steps) and leaves others open (a required tool may be in the part not recorded), so those criteria count it as missing, and the interval allows it to have gone either way.

Inspect the failures

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case case_035

The second prints one case in full. Trimmed:

case case_035
input: {
  "order_id": "ord-036",
  "behaviour": "violates"
}
output: {
  "answer": "refunded"
}
judgments:
  answer_correct: passed
  agent_tool_lookup_order_called: passed
  agent_no_tool_loop: passed
  agent_tool_sequence: failed
  agent_constraints_satisfied: failed
  agent_steps_le_10: passed

The customer got the refund (task success) from an agent that deleted a customer on the way (a constraint broken). Neither result implies the other. The full trajectory, every step with its arguments and result, is in the exported bundle (oloproof export RUN_ID) and on the case view in the workbench. The meaningful next action is in the application: stop the loop from calling a tool it must never call.

Make a candidate change and compare

In app.py, make the loop refuse the forbidden tool:

    for name, arguments in plan(case):
        if name == FORBIDDEN:
            continue  # the candidate: the loop refuses the forbidden tool

Write a comparison policy, compare.yaml:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answers-not-worse
    metric: answer_correct
    kind: non_inferiority
    margin: 0.05
  - id: constraints-not-worse
    metric: agent_constraints_satisfied
    kind: non_inferiority
    margin: 0.05
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

The candidate run alone: no-deletion now PASSes, agent_constraints_satisfied reads 40 / 40, and the gate blocks with exit 3 because the two floors are still INSUFFICIENT_EVIDENCE. The comparison:

Comparison sha256:72836e90… of run_01M4FCFH0K2RRCN3ARAEN189X7 against run_01M4FCF6544JRDB16NJ1ZFPVRZ · 40 paired cases
answer_correct: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
agent_tool_lookup_order_called: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_no_tool_loop: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_tool_sequence: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_constraints_satisfied: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_steps_le_10: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
steps_p95: +0 steps [+0, no bound] steps · p95 of per-case differences · 39 paired · 1 missing
tool_calls_p50: +0 calls [+0, +0] calls · p50 of per-case differences · 39 paired · 1 missing
72 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  constraints-not-worse  agent_constraints_satisfied  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 10 more paired cases would decide it, if the difference holds (50 in total at 8% discordance)
Gate: BLOCK (exit 3)

Read the two halves separately. On this suite, the change removed every observed deletion, which the single-run observed-count rule already settles. Whether the candidate is not worse than the baseline in general is a different question, and 40 paired cases cannot yet establish it within a 5-point margin; the planning line says roughly how many more would. No answer changed, so task success is untouched by the fix.

Part 2: a team of agents

oloproof init --example triage_agents my-team
cd my-team

The files and the recording

app.py runs three agents in one loop: triage hands each request to billing or tech, each specialist calls its own tools, and a refund billing may not issue is handed to a person. Each step names the agent that took it, and each transfer of control is a handoff step:

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

A trajectory names the agent of every step or of none; one that names only some is refused. A case declares the route it should take:

{"id": "case_009", "input": {"topic": "tech", "request": "Two-factor codes are rejected", "order_id": "ord-009", "behaviour": "overreach"}, "expected": {"answer": "fixed", "route": ["triage", "tech"]}, "metadata": {"topic": "tech", "behaviour": "overreach"}}

Evaluators and policy

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
min_slice_support: 3
  • agent_route compares the agents that held control (repeats collapsed, a hand-off's receiver included) with expected.route. A routing check, not a success check.
  • agent_tool_permissions checks every call against a closed map: an agent the map does not list may call no tool.
  • agent_max_handoffs bounds how often control changed hands.

release.yaml has answer-floor (min: 0.80), routing-floor (min: 0.70) and no-overreach, an observed-count rule with max_failures: 0 on agent_tool_permissions.

Run the team

oloproof run
Gate: BLOCK (exit 1)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum      │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-overreach  │ agent_tool_permissions │ FAIL                  │ observed_failures_exceed_limit │
│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
oloproof inspect RUN_ID --failures
6 of 30 cases failed, errored or did not finish

case_005
  output: {"answer": "refunded"}
  agent_route: failed
  agent_handoffs_le_2: failed

case_009
  output: {"answer": "fixed"}
  agent_tool_permissions: failed
...
case_030
  output: {"answer": "unresolved"}
  answer_correct: failed
  agent_route: failed
  agent_tool_permissions: error: MissingFieldError: truncated_trajectory: the trace stops before whether an agent called a tool it was not given is settled
  agent_handoffs_le_2: failed
  • case_009 and case_020: tech issued a refund, a tool only billing holds. Both answered correctly. Task success, permission broken.
  • case_005 and two others went to the wrong specialist first and came back through triage: the answer is right, the route and the hand-off bound are not.
  • case_030 bounced between billing and tech until the loop's bound. Its truncated trace already proves the route and hand-off failures, and cannot settle permissions, so that criterion is missing for it rather than passed.

Nothing in the output says which agent is to blame. A route divergence says where two routes part; that an agent caused a failure is a claim about what would have happened had it acted otherwise, which no check here makes.

Change the team and compare

The meaningful next action for the permission failures: tech hands a refund to billing instead of issuing it. In app.py, in tech:

        # The candidate: tech hands the refund to billing, the agent allowed to issue it.
        trace.hand_off("tech", "billing", "a goodwill refund")
        trace.call("billing", "issue_refund", order_id=str(case["order_id"]))

With a compare.yaml holding answers-not-worse on answer_correct and routing-not-worse on agent_route, both non_inferiority with margin: 0.05:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

The candidate run alone:

Gate: BLOCK (exit 3)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum    │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold  │
│ no-overreach  │ agent_tool_permissions │ INSUFFICIENT_EVIDENCE │ missing_could_change_outcome │
│ agent_route            │ 80.0%    │ [61.4%, 92.3%]  │ 24 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 100.0%   │ [82.7%, 100.0%] │ 29 / 29 observed · 1 missing · 0 excluded │

and the comparison:

answer_correct: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
agent_route: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
agent_tool_permissions: +6.9 points [-18.9, +33.5] · 29 paired · 1 missing · 0 excluded
agent_handoffs_le_2: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  routing-not-worse  agent_route  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    no sample size would make this PASS: the difference itself (-6.7 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

Three things to take from it:

  • No observed call broke a permission, but no-overreach is now INSUFFICIENT_EVIDENCE rather than PASS: the truncated case_030 could hide a violation in the part not recorded (missing_could_change_outcome). Fixing that loop, not the permission map, is what would settle the rule.
  • The fix changed the route of the two cases to triage > tech > billing, which their expected.route does not declare, so agent_route and agent_handoffs_le_2 fell. Whether that route is now correct is a product decision: if it is, update the cases' expected.route; a routing check measures conformance to what you declared, not quality.
  • The planning line says more cases would move routing-not-worse toward FAIL, not PASS. The comparison is telling you the candidate as written trades routing for permissions.

Troubleshooting

SymptomCauseFix
Agent evaluators refuse to runrecords: [agent_trajectory/v1] missing on the systemDeclare it in oloproof.yaml and on @system
A trajectory is refusedSome steps name an agent and others do notName the agent of every step, or of none
Many cases missing on a criterionTruncated traces: the loop hit its boundRaise the bound, or fix the loop; missing cases widen the interval rather than pass
agent_tool_sequence has a small denominatorCases without expected.toolsDeclare the sequence where it matters; [] means "expects no tool"
A constraint metric never failsThe application does not record that checkRecord an AgentConstraintCheck where your environment observes it
Results differ between runs of the same versionTools read or write shared stateReset that state before each case in your application; Oloproof does not
An agent added to the team fails permissions at onceThe permission map is closedDeclare what the new agent may call

Limitations

  • Oloproof does not drive, sandbox or reset an agent. Tool side effects, sessions, state and their reset belong to your application.
  • Every check reads the recorded trajectory. What the application does not record cannot be measured, and a truncated trace counts as missing wherever its prefix does not settle the question.
  • Trajectory checks are deterministic rules. There is no LLM-judged trajectory quality check.
  • No output attributes a failure to a step or an agent. Agent replay, which re-runs a case from a recorded checkpoint with a step dropped to label it needed or unnecessary, exists in the Python SDK only (replay_case), for a system that implements replay from its checkpoints; there is no CLI command for it, and nothing above uses it.
  • Multi-turn conversations are a different surface (SDK only); see What works today.
  • The examples' plan and behaviour fields stand in for a model's decisions so the runs are reproducible. A live model in your loop calls a provider, needs credentials and costs money per case.