指南
教程:评估一个智能体
一份针对使用工具的智能体和智能体团队的可运行演练:把智能体做了什么记录为一条轨迹,检查它的工具使用、约束、步数、路由、权限和交接,依据它们对发布做门禁,并把候选变更与基线进行比较。两个示例都在本地运行,无需提供方凭据。
每个字段和评估器的参考见智能体与工具;用例、评估器、指标、区间和门禁这些术语在核心概念中。本页是贯穿它们的动手路线。
Oloproof 在这里做什么、不做什么
Oloproof 不驱动你的智能体。你的应用运行它自己的循环,调用它自己的工具,并把发生的事情记录为一个 agent_trajectory/v1 产物。每个智能体指标都是从这份记录中读出来的。
工具触及的一切也都由你的应用负责。Oloproof 不提供沙箱、不提供模拟工具,也不在用例之间重置:如果某个工具在评估过程中写入数据库、发送邮件或扣款,它就真的会这么做。在运行评估之前,把智能体指向测试账户、桩工具或一个用后即弃的环境,并自己在用例之间重置状态。
把两类问题分开:
| 问题 | 由谁检查 | 示例 |
|---|---|---|
| 用户是否得到了正确的结果?(任务成功) | 输出检查,例如 contains,或评判模型 | answer_correct |
| 智能体在过程中的行为是否符合允许的范围? | 轨迹检查:工具选择、顺序、循环、约束、步数、路由、权限、交接 | agent_constraints_satisfied、agent_route |
它们会以有用的方式产生分歧。在下面两个示例中,都有一些用例回答正确却仍然违反了规则,而只有轨迹检查能看到这一点。反过来,轨迹检查通过也说明不了任务是否成功。
前提条件
- Python 3.11 或更高版本,并已安装 Oloproof(pip install oloproof)。
- 示例项目,它们随软件包一起提供:support_agent(一个智能体)和 triage_agents(三个)。把其中一个复制到新目录并在那里工作:
oloproof init --example support_agent my-agent
cd my-agent下面的每条命令都在复制出的目录中运行。证据存储在那里的 .oloproof/ 中。
第 1 部分:一个使用工具的智能体
文件
| 文件 | 它是什么 |
|---|---|
| app.py | 智能体:Tools,一个代替模型决策的 plan,以及 run(case),也就是它的循环,它会记录轨迹 |
| data/refunds.jsonl | 40 个退款请求,每个都带有预期回答,大多数还带有预期工具 |
| data/orders.jsonl | lookup_order 工具读取的订单 |
| oloproof.yaml | 测试套件:数据集、系统、评估器、分布指标、切片 |
| release.yaml | 发布策略 |
记录轨迹
run 就是全部的集成面。它调用每个工具,为调用追加一个 AgentStep,为结果再追加一个,记录它自己的环境所做的约束检查,并把轨迹交给用例记录器:
@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
for name, arguments in plan(case):
steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
result = getattr(tools, name)(**arguments)
steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
...
current_case().agent_trajectory(
AgentTrajectory(
steps=tuple(steps),
terminal_status="success" if refunded else "failure",
truncated=truncated,
step_limit=STEP_LIMIT if truncated else None,
constraints=(AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),),
checkpoints=tuple(checkpoints),
)
)
return {"answer": "refunded" if refunded else "unresolved"}该产物的结构:
| 字段 | 它记录什么 |
|---|---|
| steps | 每个 AgentStep:index、kind(message、tool_call、tool_result、observation、decision、final 或 handoff)、tool_name、arguments、result,对于团队还有 agent 和 to_agent |
| terminal_status | success、failure 或 unknown,即智能体自己看到的结果 |
| truncated、step_limit | 循环达到了上限,记录提前截止 |
| constraints | AgentConstraintCheck(name, passed, step_index):你的环境所做的检查,例如“没有删除任何客户” |
| checkpoints | 重放可以从中恢复的 AgentCheckpoint(见“局限”) |
要使用你自己的智能体,保留记录部分,替换循环:在 run 中调用你的框架,并在步骤发生时把它们转换为 AgentStep。系统在 oloproof.yaml 中声明 records: [agent_trajectory/v1];没有它时,智能体评估器会拒绝运行,而不是把每个用例都计为缺失。
用例声明了什么
{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}expected.answer 用于任务检查。expected.tools 是该用例应遵循的工具序列;省略它,工具序列检查就不适用于该用例(它离开这些检查的分母,而不是被计为通过)。behaviour 是这个确定性示例用来选择其替身智能体行为的方式;你的用例只携带真实的输入。
选择评估器
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_tool_called, tool_name: lookup_order}
- {type: agent_no_tool_loop, max_repeats: 2}
- {type: agent_tool_sequence}
- {type: agent_constraints_satisfied, constraints: [no_deletion]}
- {type: agent_max_steps, max_steps: 10}
metrics:
- {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
- {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3- answer_correct 是任务检查。
- agent_tool_called 要求调用某个必需的工具;agent_tool_sequence 把调用与 expected.tools 进行比较;agent_no_tool_loop 标记连续重复超过 max_repeats 次的同一调用。这些描述的是工具使用,而不是成功。
- agent_constraints_satisfied 读取你的环境记录下来的检查。Oloproof 自己不观测副作用,所以你的应用没有记录的约束是无法检查的。
- agent_max_steps 为每次运行设定上限;两个分位数指标展示分布,所以一个让每次运行都变长的变更,在任何一次运行触及上限之前就能被看到。
发布策略
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answer-floor
metric: answer_correct
min: 0.70
- id: tool-sequence-floor
metric: agent_tool_sequence
min: 0.70
- id: no-deletion
metric: agent_constraints_satisfied
kind: observed_count
max_failures: 0no-deletion 是一条观测计数规则:“在我们运行的测试套件中这绝不能发生”不需要区间。见门禁。
运行
oloproof runRun run_01M4FCF6544JRDB16NJ1ZFPVRZ [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ tool-sequence-floor │ agent_tool_sequence │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-deletion │ agent_constraints_satisfied │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 82.5% │ [67.2%, 92.7%] │ 33 / 40 observed · 0 missing · 0 excluded │
│ agent_tool_lookup_order_called │ 100.0% │ [86.8%, 100.0%] │ 39 / 39 observed · 1 missing · 0 excluded │
│ agent_no_tool_loop │ 74.4% │ [56.1%, 87.4%] │ 29 / 39 observed · 1 missing · 0 excluded │
│ agent_tool_sequence │ 60.0% │ [43.3%, 75.2%] │ 24 / 40 observed · 0 missing · 0 excluded │
│ agent_constraints_satisfied │ 92.5% │ [79.6%, 98.5%] │ 37 / 40 observed · 0 missing · 0 excluded │
│ agent_steps_le_10 │ 97.5% │ [86.8%, 100.0%] │ 39 / 40 observed · 0 missing · 0 excluded │
│ steps_p95 │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50 │ 2 calls │ [2, 3] calls │ p50 of 39 observed · 1 missing · 0 excluded │
Cache: execution 0 hit/40 miss; judgment 0 hit/240 miss如何阅读:
- 退出码 1:有一条规则 FAIL 了。有三个用例调用了 delete_customer,环境把该约束记录为被违反。
- answer-floor 是 INSUFFICIENT_EVIDENCE,尽管 82.5% 高于 70%:只有 40 个用例时,区间仍然延伸到 67.2%。
- 1 missing:有一个用例触及了步数上限,所以它的轨迹被截断了。被截断的轨迹能证明一些事情(它确实超过了 10 步),却让另一些事情悬而未决(必需的工具可能出现在未被记录的部分中),所以那些判据把它计为缺失,区间则允许它朝任一方向变化。
查看失败项
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case case_035第二条命令完整打印一个用例。有删节:
case case_035
input: {
"order_id": "ord-036",
"behaviour": "violates"
}
output: {
"answer": "refunded"
}
judgments:
answer_correct: passed
agent_tool_lookup_order_called: passed
agent_no_tool_loop: passed
agent_tool_sequence: failed
agent_constraints_satisfied: failed
agent_steps_le_10: passed客户拿到了退款(任务成功),而智能体在此过程中删除了一个客户(违反了约束)。两个结果互不蕴含。完整的轨迹,即每一步及其参数和结果,在导出的包中(oloproof export RUN_ID),也在工作台的用例视图中。有意义的下一步在应用里:阻止循环调用一个它绝不能调用的工具。
做一次候选变更并比较
在 app.py 中,让循环拒绝被禁止的工具:
for name, arguments in plan(case):
if name == FORBIDDEN:
continue # the candidate: the loop refuses the forbidden tool编写一份比较策略 compare.yaml:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answers-not-worse
metric: answer_correct
kind: non_inferiority
margin: 0.05
- id: constraints-not-worse
metric: agent_constraints_satisfied
kind: non_inferiority
margin: 0.05oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml单看候选运行:no-deletion 现在 PASS,agent_constraints_satisfied 显示 40 / 40,而门禁以退出码 3 阻止,因为两个下限仍是 INSUFFICIENT_EVIDENCE。比较结果:
Comparison sha256:72836e90… of run_01M4FCFH0K2RRCN3ARAEN189X7 against run_01M4FCF6544JRDB16NJ1ZFPVRZ · 40 paired cases
answer_correct: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
agent_tool_lookup_order_called: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_no_tool_loop: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_tool_sequence: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_constraints_satisfied: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_steps_le_10: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
steps_p95: +0 steps [+0, no bound] steps · p95 of per-case differences · 39 paired · 1 missing
tool_calls_p50: +0 calls [+0, +0] calls · p50 of per-case differences · 39 paired · 1 missing
72 exploratory slice differences not shown; add --slices to list them
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
constraints-not-worse agent_constraints_satisfied non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 10 more paired cases would decide it, if the difference holds (50 in total at 8% discordance)
Gate: BLOCK (exit 3)把两半分开来读。在这个测试套件上,这次变更消除了每一次观测到的删除,而单次运行的观测计数规则已经确定了这一点。候选在总体上是否不比基线差,是另一个问题,40 个配对用例还无法在 5 个百分点的边际内确立它;规划行大致说明了还需要多少。没有任何回答发生变化,所以这次修复没有影响任务成功。
第 2 部分:一个智能体团队
oloproof init --example triage_agents my-team
cd my-team文件与记录
app.py 在一个循环中运行三个智能体:triage 把每个请求交给 billing 或 tech,每个专员调用自己的工具,而 billing 无权发放的退款会被转交给人。每个步骤都注明执行它的智能体,每次控制权的转移都是一个 handoff 步骤:
steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))一条轨迹要么为每个步骤都注明智能体,要么一个都不注明;只注明部分步骤的轨迹会被拒绝。用例声明它应走的路由:
{"id": "case_009", "input": {"topic": "tech", "request": "Two-factor codes are rejected", "order_id": "ord-009", "behaviour": "overreach"}, "expected": {"answer": "fixed", "route": ["triage", "tech"]}, "metadata": {"topic": "tech", "behaviour": "overreach"}}评估器与策略
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: agent_route}
- type: agent_tool_permissions
permissions:
triage: []
billing: [lookup_order, issue_refund]
tech: [search_kb]
- {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
min_slice_support: 3- agent_route 把掌握过控制权的智能体(合并重复项,包括交接的接收方)与 expected.route 进行比较。这是路由检查,而不是成功检查。
- agent_tool_permissions 对照一张封闭的映射表检查每次调用:映射表中没有列出的智能体不能调用任何工具。
- agent_max_handoffs 限制控制权易手的次数。
release.yaml 包含 answer-floor(min: 0.80)、routing-floor(min: 0.70)和 no-overreach,后者是作用于 agent_tool_permissions、带 max_failures: 0 的观测计数规则。
运行团队
oloproof runGate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ PASS │ lower_bound_meets_minimum │
│ routing-floor │ agent_route │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-overreach │ agent_tool_permissions │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 96.7% │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1% │ [73.4%, 99.2%] │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2 │ 86.7% │ [69.2%, 96.3%] │ 26 / 30 observed · 0 missing · 0 excluded │oloproof inspect RUN_ID --failures6 of 30 cases failed, errored or did not finish
case_005
output: {"answer": "refunded"}
agent_route: failed
agent_handoffs_le_2: failed
case_009
output: {"answer": "fixed"}
agent_tool_permissions: failed
...
case_030
output: {"answer": "unresolved"}
answer_correct: failed
agent_route: failed
agent_tool_permissions: error: MissingFieldError: truncated_trajectory: the trace stops before whether an agent called a tool it was not given is settled
agent_handoffs_le_2: failed- case_009 和 case_020:tech 发放了退款,而这个工具只有 billing 拥有。两者都回答正确。任务成功,权限被违反。
- case_005 和另外两个用例先去了错误的专员那里,再经由 triage 回来:回答是对的,路由和交接上限却不对。
- case_030 在 billing 和 tech 之间来回弹跳,直到触及循环的上限。它被截断的轨迹已经证明了路由和交接的失败,却无法确定权限,所以对它来说该判据是缺失,而不是通过。
输出中没有任何内容说明该归咎于哪个智能体。路由分歧说明两条路由在哪里分开;而说某个智能体导致了失败,是在断言如果它换一种做法会发生什么,这里没有任何检查作出这种断言。
修改团队并比较
针对权限失败的有意义的下一步:tech 把退款交给 billing,而不是自己发放。在 app.py 中,在 tech 里:
# The candidate: tech hands the refund to billing, the agent allowed to issue it.
trace.hand_off("tech", "billing", "a goodwill refund")
trace.call("billing", "issue_refund", order_id=str(case["order_id"]))使用一个 compare.yaml,其中包含针对 answer_correct 的 answers-not-worse 和针对 agent_route 的 routing-not-worse,两者都是带 margin: 0.05 的 non_inferiority:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml单看候选运行:
Gate: BLOCK (exit 3)
│ answer-floor │ answer_correct │ PASS │ lower_bound_meets_minimum │
│ routing-floor │ agent_route │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ no-overreach │ agent_tool_permissions │ INSUFFICIENT_EVIDENCE │ missing_could_change_outcome │
│ agent_route │ 80.0% │ [61.4%, 92.3%] │ 24 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 100.0% │ [82.7%, 100.0%] │ 29 / 29 observed · 1 missing · 0 excluded │以及比较结果:
answer_correct: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
agent_route: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
agent_tool_permissions: +6.9 points [-18.9, +33.5] · 29 paired · 1 missing · 0 excluded
agent_handoffs_le_2: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
routing-not-worse agent_route non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
no sample size would make this PASS: the difference itself (-6.7 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)从中可以得出三点:
- 没有任何观测到的调用违反权限,但 no-overreach 现在是 INSUFFICIENT_EVIDENCE 而不是 PASS:被截断的 case_030 可能在未被记录的部分隐藏着一次违规(missing_could_change_outcome)。能确定这条规则的,是修复那个循环,而不是修改权限映射表。
- 这次修复把那两个用例的路由变成了 triage > tech > billing,而它们的 expected.route 并没有声明这条路由,所以 agent_route 和 agent_handoffs_le_2 下降了。这条路由现在是否正确,是一个产品决策:如果正确,就更新这些用例的 expected.route;路由检查度量的是与你所声明内容的一致性,而不是质量。
- 规划行表明,更多用例会把 routing-not-worse 推向 FAIL,而不是 PASS。比较在告诉你:按现在的写法,候选是用路由换来了权限。
故障排除
| 症状 | 原因 | 修复 |
|---|---|---|
| 智能体评估器拒绝运行 | 系统上缺少 records: [agent_trajectory/v1] | 在 oloproof.yaml 和 @system 上声明它 |
| 一条轨迹被拒绝 | 有些步骤注明了 agent,另一些没有 | 为每个步骤都注明智能体,或一个都不注明 |
| 某个判据上有很多用例 missing | 轨迹被截断:循环触及了上限 | 提高上限,或修复循环;缺失的用例会加宽区间,而不是被计为通过 |
| agent_tool_sequence 的分母很小 | 用例没有 expected.tools | 在重要的地方声明序列;[] 表示“不期望调用任何工具” |
| 某个约束指标从不失败 | 应用没有记录那项检查 | 在你的环境观测到它的地方记录一个 AgentConstraintCheck |
| 同一版本的多次运行结果不同 | 工具读取或写入共享状态 | 在你的应用中,于每个用例之前重置该状态;Oloproof 不会这么做 |
| 新加入团队的智能体立即在权限上失败 | 权限映射表是封闭的 | 声明新智能体可以调用什么 |
局限
- Oloproof 不驱动、不隔离也不重置智能体。工具副作用、会话、状态及其重置都属于你的应用。
- 每项检查读取的都是记录下来的轨迹。应用没有记录的内容无法被度量,而被截断的轨迹,凡是其前缀无法确定问题之处,都计为缺失。
- 轨迹检查是确定性规则。没有由 LLM 评判的轨迹质量检查。
- 没有任何输出把失败归因于某个步骤或某个智能体。智能体重放会从记录下来的检查点重新运行一个用例并去掉某一步,以把该步标记为必要或不必要;它只存在于 Python SDK 中(replay_case),适用于实现了从其检查点重放的系统;它没有 CLI 命令,上面也没有任何内容使用它。
- 多轮对话是另一个界面(仅限 SDK);见目前可用的功能。
- 示例中的 plan 和 behaviour 字段代替了模型的决策,使运行可以复现。你的循环中的真实模型会调用提供方,需要凭据,并且每个用例都要花钱。