本文へスキップ

ガイド

チュートリアル: エージェントを評価する

ツールを使う 1 つのエージェントと、エージェントのチームを実際に動かして進む手順です。エージェントが行ったことを軌跡として記録し、ツールの使い方、制約、ステップ数、ルーティング、権限、ハンドオフを確認し、それらに基づいてリリースをゲートし、候補の変更をベースラインと比較します。どちらの例もプロバイダーの認証情報なしにローカルで実行できます。

すべてのフィールドと評価器のリファレンスはエージェントとツールにあります。ケース、評価器、メトリクス、区間、ゲートという用語は基本概念にあります。このページは、それらを手を動かしながらたどる道筋です。

ここで Oloproof が行うこと、行わないこと

Oloproof はあなたのエージェントを動かしません。あなたのアプリケーションが自身のループを回し、自身のツールを呼び出し、起きたことを agent_trajectory/v1 アーティファクトとして記録します。エージェントのメトリクスはすべてその記録から読み取られます。

ツールが触れるものもすべてあなたのアプリケーションが持ちます。Oloproof はサンドボックスも、シミュレートされたツールも、ケース間のリセットも提供しません。評価中にツールがデータベースに書き込んだり、メールを送ったり、カードに請求したりすれば、それは実際に起こります。評価を実行する前に、エージェントをテスト用アカウント、スタブ化したツール、使い捨ての環境に向け、ケース間の状態は自分でリセットしてください。

2 種類の問いを分けて考えてください。

問い確認するもの例
ユーザーは正しい結果を得たか?(タスクの成功)contains のような出力の確認、またはジャッジanswer_correct
エージェントは途中で許された振る舞いをしたか?軌跡の確認: ツールの選択、順序、ループ、制約、ステップ数、ルーティング、権限、ハンドオフagent_constraints_satisfied、agent_route

この 2 つは有益な形で食い違います。以下の両方の例で、正しく答えながらルールを破るケースがあり、それに気づくのは軌跡の確認だけです。軌跡の確認が合格しても、タスクが成功したかどうかについては何もわかりません。

前提条件

  • Python 3.11 以降と、インストール済みの Oloproof(pip install oloproof)。
  • パッケージに同梱されているサンプルプロジェクト: support_agent(1 つのエージェント)と triage_agents(3 つ)。いずれかを新しいディレクトリにコピーし、そこで作業してください。
oloproof init --example support_agent my-agent
cd my-agent

以下のコマンドはすべて、コピーしたディレクトリの中で実行します。エビデンスはそこにある .oloproof/ に保存されます。

パート 1: ツールを使う 1 つのエージェント

ファイル

ファイル内容
app.pyエージェント: Tools、モデルの判断の代わりとなる plan、そしてそのループであり軌跡を記録する run(case)
data/refunds.jsonl40 件の返金リクエスト。それぞれに期待される回答があり、ほとんどには期待されるツールもあります
data/orders.jsonllookup_order ツールが読む注文
oloproof.yamlスイート: データセット、システム、評価器、分布のメトリクス、スライス
release.yamlリリースポリシー

軌跡の記録

run が統合面のすべてです。各ツールを呼び出し、呼び出しについて 1 つ、その結果について 1 つの AgentStep を追加し、自身の環境が行った制約の確認を記録して、軌跡をケースレコーダーに渡します。

@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
    ...
    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),),
            checkpoints=tuple(checkpoints),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

アーティファクトの形は次のとおりです。

フィールド記録するもの
steps各 AgentStep: index、kind(message、tool_call、tool_result、observation、decision、final、handoff)、tool_name、arguments、result、そしてチームの場合は agent と to_agent
terminal_statusエージェントから見た success、failure、unknown
truncated、step_limitループが上限に達し、記録が途中で終わっていること
constraintsAgentConstraintCheck(name, passed, step_index): 「顧客は削除されなかった」のような、環境が行った確認
checkpointsリプレイが再開できる AgentCheckpoint(制限事項を参照)

自分のエージェントを使うには、記録はそのままにしてループを置き換えます。run の中で自分のフレームワークを呼び出し、そのステップを起きるたびに AgentStep に変換してください。システムは oloproof.yaml で records: [agent_trajectory/v1] を宣言します。これがなければ、エージェントの評価器はすべてのケースを欠測として数えるのではなく、実行を拒否します。

ケースが宣言するもの

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.answer はタスクの確認のためのものです。expected.tools はケースがたどるべきツールの順序です。省略すると、ツール順序の確認はそのケースに適用されません(合格するのではなく、分母から外れます)。behaviour は、この決定的な例が代役のエージェントの行動を選ぶ方法です。あなたのケースは実際の入力だけを持ちます。

評価器の選択

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
  • answer_correct はタスクの確認です。
  • agent_tool_called は必須のツールを求めます。agent_tool_sequence は呼び出しを expected.tools と比較します。agent_no_tool_loop は同じ呼び出しが max_repeats 回を超えて連続したことを示します。これらはツールの使い方を表すものであり、成功を表すものではありません。
  • agent_constraints_satisfied は環境が記録した確認を読みます。Oloproof 自身は副作用を観測しないため、アプリケーションが記録しない制約は確認できません。
  • agent_max_steps は各実行を上限で抑えます。2 つの分位点メトリクスは分布を示すため、すべての実行を長くする変更は、どれか 1 つの実行が上限に達する前に見えるようになります。

リリースポリシー

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: tool-sequence-floor
    metric: agent_tool_sequence
    min: 0.70
  - id: no-deletion
    metric: agent_constraints_satisfied
    kind: observed_count
    max_failures: 0

no-deletion は観測件数ルールです。「実行したスイートでこれが起きてはならない」には区間は必要ありません。ゲーティングを参照してください。

実行する

oloproof run
Run run_01M4FCF6544JRDB16NJ1ZFPVRZ [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor        │ answer_correct              │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ tool-sequence-floor │ agent_tool_sequence         │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
Cache: execution 0 hit/40 miss; judgment 0 hit/240 miss

読み方は次のとおりです。

  • 終了コード 1: ルールが FAIL しました。3 つのケースが delete_customer を呼び出し、環境はその制約が破られたと記録しました。
  • answer-floor は 82.5% が 70% を上回っているにもかかわらず INSUFFICIENT_EVIDENCE です。40 ケースでは、区間はまだ 67.2% まで届きます。
  • 1 missing: 1 つのケースがステップの上限に達したため、そのトレースは途中で切れています。切れたトレースはあることを証明し(10 ステップを超えたことは確か)、別のことは未決のままにします(必須のツールが記録されていない部分にあるかもしれない)。そのため、それらの基準はそのケースを欠測として数え、区間はどちらの結果もありえたものとして扱います。

失敗を調べる

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case case_035

2 つ目は 1 つのケースを完全に出力します。抜粋は次のとおりです。

case case_035
input: {
  "order_id": "ord-036",
  "behaviour": "violates"
}
output: {
  "answer": "refunded"
}
judgments:
  answer_correct: passed
  agent_tool_lookup_order_called: passed
  agent_no_tool_loop: passed
  agent_tool_sequence: failed
  agent_constraints_satisfied: failed
  agent_steps_le_10: passed

顧客は返金を受けました(タスクの成功)が、エージェントは途中で顧客を削除しました(制約違反)。どちらの結果も他方を意味しません。すべてのステップとその引数および結果を含む完全な軌跡は、エクスポートしたバンドル(oloproof export RUN_ID)とワークベンチのケース画面にあります。意味のある次の行動はアプリケーションの中にあります。決して呼んではならないツールをループが呼べないようにすることです。

候補の変更を加えて比較する

app.py で、禁止されたツールをループが拒否するようにします。

    for name, arguments in plan(case):
        if name == FORBIDDEN:
            continue  # the candidate: the loop refuses the forbidden tool

比較ポリシー compare.yaml を書きます。

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answers-not-worse
    metric: answer_correct
    kind: non_inferiority
    margin: 0.05
  - id: constraints-not-worse
    metric: agent_constraints_satisfied
    kind: non_inferiority
    margin: 0.05
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

候補の実行単独では、no-deletion は PASS になり、agent_constraints_satisfied は 40 / 40 となり、2 つの下限ルールがまだ INSUFFICIENT_EVIDENCE であるため、ゲートは終了コード 3 でブロックします。比較は次のとおりです。

Comparison sha256:72836e90… of run_01M4FCFH0K2RRCN3ARAEN189X7 against run_01M4FCF6544JRDB16NJ1ZFPVRZ · 40 paired cases
answer_correct: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
agent_tool_lookup_order_called: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_no_tool_loop: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_tool_sequence: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_constraints_satisfied: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_steps_le_10: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
steps_p95: +0 steps [+0, no bound] steps · p95 of per-case differences · 39 paired · 1 missing
tool_calls_p50: +0 calls [+0, +0] calls · p50 of per-case differences · 39 paired · 1 missing
72 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  constraints-not-worse  agent_constraints_satisfied  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 10 more paired cases would decide it, if the difference holds (50 in total at 8% discordance)
Gate: BLOCK (exit 3)

2 つの部分を分けて読んでください。このスイートでは、変更は観測されたすべての削除をなくしました。それは単一実行の観測件数ルールがすでに決着させています。候補が一般的にベースラインより悪くないかどうかは別の問いであり、40 の対応のあるケースでは、まだ 5 ポイントのマージン内でそれを確立できません。計画の行は、あとおよそ何件あれば確立できるかを示しています。どの回答も変わらなかったため、タスクの成功はこの修正の影響を受けていません。

パート 2: エージェントのチーム

oloproof init --example triage_agents my-team
cd my-team

ファイルと記録

app.py は 1 つのループで 3 つのエージェントを動かします。triage は各リクエストを billing か tech に渡し、各専門エージェントは自身のツールを呼び出し、billing が発行してはならない返金は人に渡されます。各ステップはそれを行ったエージェントを示し、制御の移転はそれぞれ handoff ステップです。

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

軌跡は、すべてのステップのエージェントを示すか、どれも示さないかのどちらかです。一部だけを示す軌跡は拒否されます。ケースはたどるべきルートを宣言します。

{"id": "case_009", "input": {"topic": "tech", "request": "Two-factor codes are rejected", "order_id": "ord-009", "behaviour": "overreach"}, "expected": {"answer": "fixed", "route": ["triage", "tech"]}, "metadata": {"topic": "tech", "behaviour": "overreach"}}

評価器とポリシー

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
min_slice_support: 3
  • agent_route は、制御を持ったエージェント(繰り返しはまとめ、ハンドオフの受け手を含む)を expected.route と比較します。これはルーティングの確認であり、成功の確認ではありません。
  • agent_tool_permissions はすべての呼び出しを閉じたマップに照らして確認します。マップに載っていないエージェントはどのツールも呼べません。
  • agent_max_handoffs は、制御が持ち替えられた回数を上限で抑えます。

release.yaml には answer-floor(min: 0.80)、routing-floor(min: 0.70)、そして agent_tool_permissions に対する max_failures: 0 の観測件数ルール no-overreach があります。

チームを実行する

oloproof run
Gate: BLOCK (exit 1)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum      │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-overreach  │ agent_tool_permissions │ FAIL                  │ observed_failures_exceed_limit │
│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
oloproof inspect RUN_ID --failures
6 of 30 cases failed, errored or did not finish

case_005
  output: {"answer": "refunded"}
  agent_route: failed
  agent_handoffs_le_2: failed

case_009
  output: {"answer": "fixed"}
  agent_tool_permissions: failed
...
case_030
  output: {"answer": "unresolved"}
  answer_correct: failed
  agent_route: failed
  agent_tool_permissions: error: MissingFieldError: truncated_trajectory: the trace stops before whether an agent called a tool it was not given is settled
  agent_handoffs_le_2: failed
  • case_009 と case_020: tech が返金を発行しましたが、それは billing だけが持つツールです。どちらも正しく答えました。タスクは成功、権限は違反です。
  • case_005 とほかの 2 つは、最初に誤った専門エージェントに行き、triage を通って戻ってきました。回答は正しいですが、ルートとハンドオフの上限はそうではありません。
  • case_030 はループの上限まで billing と tech の間を行き来しました。その切れたトレースはルートとハンドオフの失敗をすでに証明していますが、権限については決着できないため、その基準ではそのケースは合格ではなく欠測です。

出力のどこにも、どのエージェントに責任があるかは書かれていません。ルートの分岐は 2 つのルートがどこで分かれるかを示します。あるエージェントが失敗を引き起こしたというのは、そのエージェントが別の行動をとっていたら何が起きたかについての主張であり、ここでの確認はどれもそれを行いません。

チームを変更して比較する

権限の失敗に対する意味のある次の行動: tech は返金を発行する代わりに billing に渡します。app.py の tech で次のようにします。

        # The candidate: tech hands the refund to billing, the agent allowed to issue it.
        trace.hand_off("tech", "billing", "a goodwill refund")
        trace.call("billing", "issue_refund", order_id=str(case["order_id"]))

answer_correct に対する answers-not-worse と agent_route に対する routing-not-worse を、どちらも margin: 0.05 の non_inferiority として持つ compare.yaml を使います。

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

候補の実行単独では次のとおりです。

Gate: BLOCK (exit 3)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum    │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold  │
│ no-overreach  │ agent_tool_permissions │ INSUFFICIENT_EVIDENCE │ missing_could_change_outcome │
│ agent_route            │ 80.0%    │ [61.4%, 92.3%]  │ 24 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 100.0%   │ [82.7%, 100.0%] │ 29 / 29 observed · 1 missing · 0 excluded │

そして比較は次のとおりです。

answer_correct: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
agent_route: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
agent_tool_permissions: +6.9 points [-18.9, +33.5] · 29 paired · 1 missing · 0 excluded
agent_handoffs_le_2: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  routing-not-worse  agent_route  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    no sample size would make this PASS: the difference itself (-6.7 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

ここから読み取るべきことが 3 つあります。

  • 観測された呼び出しで権限を破ったものはありませんが、no-overreach は PASS ではなく INSUFFICIENT_EVIDENCE になりました。切れた case_030 は、記録されていない部分に違反を隠している可能性があります(missing_could_change_outcome)。このルールを決着させるのは、権限マップではなく、そのループを直すことです。
  • 修正によって 2 つのケースのルートが triage > tech > billing に変わりましたが、それはそれらの expected.route が宣言していないものなので、agent_route と agent_handoffs_le_2 は下がりました。そのルートが今や正しいかどうかはプロダクトの判断です。正しいならケースの expected.route を更新してください。ルーティングの確認が測るのは宣言したものへの適合であり、品質ではありません。
  • 計画の行は、ケースを増やすと routing-not-worse は PASS ではなく FAIL のほうへ動くと示しています。比較は、書かれたままの候補がルーティングと引き換えに権限を改善していると伝えています。

トラブルシューティング

症状原因対処
エージェントの評価器が実行を拒否するシステムに records: [agent_trajectory/v1] がないoloproof.yaml と @system で宣言する
軌跡が拒否される一部のステップは agent を示し、他は示さないすべてのステップのエージェントを示すか、どれも示さない
ある基準で多くのケースが missing になる切れたトレース: ループが上限に達した上限を上げるか、ループを直す。欠測ケースは合格せず、区間を広げる
agent_tool_sequence の分母が小さいexpected.tools のないケース重要なところで順序を宣言する。[] は「ツールを期待しない」を意味する
制約のメトリクスが決して失敗しないアプリケーションがその確認を記録していない環境がそれを観測する場所で AgentConstraintCheck を記録する
同じバージョンの実行間で結果が異なるツールが共有状態を読み書きしているアプリケーションで各ケースの前にその状態をリセットする。Oloproof は行わない
チームに加えたエージェントがすぐに権限で失敗する権限マップは閉じている新しいエージェントが呼べるものを宣言する

制限事項

  • Oloproof はエージェントを動かさず、サンドボックス化もリセットもしません。ツールの副作用、セッション、状態とそのリセットはあなたのアプリケーションに属します。
  • すべての確認は記録された軌跡を読みます。アプリケーションが記録しないものは測定できず、切れたトレースは、その前半部分で問いが決着しない限り欠測として数えられます。
  • 軌跡の確認は決定的なルールです。LLM で判定する軌跡の品質の確認はありません。
  • 失敗をステップやエージェントに帰属させる出力はありません。エージェントのリプレイは、記録されたチェックポイントからステップを 1 つ取り除いてケースを再実行し、そのステップが必要か不要かをラベル付けするもので、Python SDK にのみ(replay_case)、チェックポイントからのリプレイを実装したシステム向けに存在します。CLI コマンドはなく、上記のどこでも使っていません。
  • 複数ターンの会話は別の面です(SDK のみ)。現在動作するものを参照してください。
  • 例の plan と behaviour フィールドは、実行を再現可能にするためにモデルの判断の代わりをしています。ループの中の実際のモデルはプロバイダーを呼び出し、認証情報が必要で、ケースごとに費用がかかります。