본문으로 건너뛰기

가이드

튜토리얼: 에이전트 평가하기

도구를 쓰는 에이전트와 에이전트 팀을 위한 실행 가능한 안내입니다. 에이전트가 한 일을 궤적으로 기록하고, 도구 사용, 제약, 단계, 라우팅, 권한, 핸드오프를 검사하며, 그것으로 릴리스를 게이트하고, 후보 변경을 기준선과 비교합니다. 두 예제 모두 제공자 자격 증명 없이 로컬에서 실행됩니다.

모든 필드와 평가기의 참조는 에이전트와 도구에 있고, 케이스, 평가기, 지표, 구간, 게이트라는 용어는 핵심 개념에 있습니다. 이 페이지는 그것들을 직접 따라가 보는 경로입니다.

여기서 Oloproof가 하는 일과 하지 않는 일

Oloproof는 에이전트를 구동하지 않습니다. 애플리케이션이 자신의 루프를 실행하고, 자신의 도구를 호출하며, 일어난 일을 agent_trajectory/v1 아티팩트로 기록합니다. 모든 에이전트 지표는 그 기록에서 읽어 냅니다.

도구가 건드리는 모든 것도 애플리케이션이 소유합니다. Oloproof는 샌드박스도, 시뮬레이션된 도구도, 케이스 사이의 초기화도 제공하지 않습니다. 평가 중에 도구가 데이터베이스에 쓰거나, 이메일을 보내거나, 카드에 결제하면 실제로 그 일이 일어납니다. 평가를 실행하기 전에 에이전트가 테스트 계정, 스텁 도구, 또는 일회용 환경을 가리키게 하고, 케이스 사이의 상태는 직접 초기화하세요.

두 종류의 질문을 구분하세요.

질문검사하는 것예
사용자가 올바른 결과를 얻었는가?(작업 성공)contains 같은 출력 검사, 또는 심사 모델answer_correct
에이전트가 그 과정에서 허용된 대로 행동했는가?궤적 검사: 도구 선택, 순서, 루프, 제약, 단계, 라우팅, 권한, 핸드오프agent_constraints_satisfied, agent_route

둘은 유용한 방식으로 엇갈립니다. 아래 두 예제 모두에서 일부 케이스는 올바르게 답하면서도 규칙을 어기며, 그것은 궤적 검사만 볼 수 있습니다. 궤적 검사를 통과했다고 해서 작업이 성공했다는 뜻도 아닙니다.

사전 준비

  • Python 3.11 이상, 그리고 Oloproof 설치(pip install oloproof).
  • 패키지와 함께 제공되는 예제 프로젝트: support_agent(에이전트 하나)와 triage_agents(셋). 하나를 새 디렉터리에 복사하고 그곳에서 작업하세요.
oloproof init --example support_agent my-agent
cd my-agent

아래의 모든 명령은 복사한 디렉터리 안에서 실행합니다. 증거는 그곳의 .oloproof/에 저장됩니다.

1부: 도구를 쓰는 에이전트 하나

파일

파일의미
app.py에이전트: Tools, 모델의 결정을 대신하는 plan, 그리고 궤적을 기록하는 루프인 run(case)
data/refunds.jsonl환불 요청 40개, 각각 기대 답변과, 대부분은 기대 도구를 가짐
data/orders.jsonllookup_order 도구가 읽는 주문
oloproof.yaml스위트: 데이터셋, 시스템, 평가기, 분포 지표, 슬라이스
release.yaml릴리스 정책

궤적 기록하기

run이 통합 지점의 전부입니다. 각 도구를 호출하고, 호출에 대한 AgentStep 하나와 결과에 대한 것 하나를 덧붙이고, 자신의 환경이 수행한 제약 검사를 기록하고, 궤적을 케이스 기록기에 넘깁니다.

@system(name="support-agent", version="slice-e-example", records=("agent_trajectory/v1",))
def run(case):
    for name, arguments in plan(case):
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_call", tool_name=name, arguments=arguments))
        result = getattr(tools, name)(**arguments)
        steps.append(AgentStep(index=len(steps) + 1, kind="tool_result", tool_name=name, result=result))
    ...
    current_case().agent_trajectory(
        AgentTrajectory(
            steps=tuple(steps),
            terminal_status="success" if refunded else "failure",
            truncated=truncated,
            step_limit=STEP_LIMIT if truncated else None,
            constraints=(AgentConstraintCheck(name="no_deletion", passed=deletion is None, step_index=...),),
            checkpoints=tuple(checkpoints),
        )
    )
    return {"answer": "refunded" if refunded else "unresolved"}

아티팩트의 형태는 다음과 같습니다.

필드기록하는 것
steps각 AgentStep: index, kind(message, tool_call, tool_result, observation, decision, final 또는 handoff), tool_name, arguments, result, 그리고 팀의 경우 agent와 to_agent
terminal_status에이전트가 본 대로 success, failure 또는 unknown
truncated, step_limit루프가 한계에 도달해 기록이 중간에 멈췄다는 것
constraintsAgentConstraintCheck(name, passed, step_index): "고객이 삭제되지 않았다" 같은, 환경이 수행한 검사
checkpoints재생이 이어서 시작할 수 있는 AgentCheckpoint(제한 사항 참조)

자신의 에이전트를 쓰려면 기록은 유지하고 루프를 바꾸세요. run에서 프레임워크를 호출하고, 그 단계를 일어나는 대로 AgentStep으로 옮기면 됩니다. 시스템은 oloproof.yaml에서 records: [agent_trajectory/v1]을 선언합니다. 없으면 에이전트 평가기는 모든 케이스를 누락으로 세는 대신 실행을 거부합니다.

케이스가 선언하는 것

{"id": "case_001", "input": {"order_id": "ord-002", "behaviour": "clean"}, "expected": {"answer": "refunded", "tools": ["lookup_order", "issue_refund"]}, "metadata": {"surface": "chat", "behaviour": "clean"}}

expected.answer는 작업 검사용입니다. expected.tools는 케이스가 따라야 할 도구 순서이며, 생략하면 도구 순서 검사가 그 케이스에 적용되지 않습니다(통과하는 것이 아니라 분모에서 빠집니다). behaviour는 이 결정론적 예제가 대역 에이전트의 행동을 고르는 방법이며, 여러분의 케이스에는 실제 입력만 담깁니다.

평가기 고르기

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_tool_called, tool_name: lookup_order}
  - {type: agent_no_tool_loop, max_repeats: 2}
  - {type: agent_tool_sequence}
  - {type: agent_constraints_satisfied, constraints: [no_deletion]}
  - {type: agent_max_steps, max_steps: 10}
metrics:
  - {id: steps_p95, type: quantile, source: agent_steps, quantile: 0.95}
  - {id: tool_calls_p50, type: quantile, source: agent_tool_calls, quantile: 0.5}
slices: [metadata.surface, first_tool, repeated_action, "trajectory_length:4,8"]
min_slice_support: 3
  • answer_correct는 작업 검사입니다.
  • agent_tool_called는 필요한 도구를 요구하고, agent_tool_sequence는 호출을 expected.tools와 비교하며, agent_no_tool_loop는 같은 호출이 연달아 max_repeats번보다 많이 반복되면 표시합니다. 이것들은 성공이 아니라 도구 사용을 설명합니다.
  • agent_constraints_satisfied는 환경이 기록한 검사를 읽습니다. Oloproof는 부수 효과를 직접 관찰하지 않으므로, 애플리케이션이 기록하지 않은 제약은 검사할 수 없습니다.
  • agent_max_steps는 각 실행을 제한하며, 두 분위수 지표는 분포를 보여 주므로, 모든 실행을 길게 만드는 변경은 어느 하나가 한계에 닿기 전에 드러납니다.

릴리스 정책

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: tool-sequence-floor
    metric: agent_tool_sequence
    min: 0.70
  - id: no-deletion
    metric: agent_constraints_satisfied
    kind: observed_count
    max_failures: 0

no-deletion은 관측 개수 규칙입니다. "우리가 실행한 스위트에서 이 일이 일어나서는 안 된다"에는 구간이 필요 없습니다. 게이팅을 보세요.

실행하기

oloproof run
Run run_01M4FCF6544JRDB16NJ1ZFPVRZ [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor        │ answer_correct              │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ tool-sequence-floor │ agent_tool_sequence         │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-deletion         │ agent_constraints_satisfied │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct                 │ 82.5%    │ [67.2%, 92.7%]       │ 33 / 40 observed · 0 missing · 0 excluded   │
│ agent_tool_lookup_order_called │ 100.0%   │ [86.8%, 100.0%]      │ 39 / 39 observed · 1 missing · 0 excluded   │
│ agent_no_tool_loop             │ 74.4%    │ [56.1%, 87.4%]       │ 29 / 39 observed · 1 missing · 0 excluded   │
│ agent_tool_sequence            │ 60.0%    │ [43.3%, 75.2%]       │ 24 / 40 observed · 0 missing · 0 excluded   │
│ agent_constraints_satisfied    │ 92.5%    │ [79.6%, 98.5%]       │ 37 / 40 observed · 0 missing · 0 excluded   │
│ agent_steps_le_10              │ 97.5%    │ [86.8%, 100.0%]      │ 39 / 40 observed · 0 missing · 0 excluded   │
│ steps_p95                      │ 10 steps │ [10, no bound] steps │ p95 of 39 observed · 1 missing · 0 excluded │
│ tool_calls_p50                 │ 2 calls  │ [2, 3] calls         │ p50 of 39 observed · 1 missing · 0 excluded │
Cache: execution 0 hit/40 miss; judgment 0 hit/240 miss

읽는 방법은 다음과 같습니다.

  • 종료 코드 1: 규칙이 FAIL했습니다. 세 케이스가 delete_customer를 호출했고, 환경이 그 제약을 위반으로 기록했습니다.
  • answer-floor는 82.5%가 70%보다 높은데도 INSUFFICIENT_EVIDENCE입니다. 케이스 40개로는 구간이 여전히 67.2%까지 이릅니다.
  • 1 missing: 한 케이스가 단계 한계에 닿아 트레이스가 잘렸습니다. 잘린 트레이스는 어떤 것은 증명하고(10단계를 넘었다는 것) 다른 것은 열어 둡니다(필요한 도구가 기록되지 않은 부분에 있을 수 있음). 그래서 해당 기준은 그것을 누락으로 세고, 구간은 어느 쪽이었을 가능성도 허용합니다.

실패 살펴보기

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case case_035

두 번째 명령은 케이스 하나를 전부 출력합니다. 줄이면 다음과 같습니다.

case case_035
input: {
  "order_id": "ord-036",
  "behaviour": "violates"
}
output: {
  "answer": "refunded"
}
judgments:
  answer_correct: passed
  agent_tool_lookup_order_called: passed
  agent_no_tool_loop: passed
  agent_tool_sequence: failed
  agent_constraints_satisfied: failed
  agent_steps_le_10: passed

고객은 환불을 받았지만(작업 성공), 에이전트는 그 과정에서 고객 하나를 삭제했습니다(제약 위반). 어느 결과도 다른 결과를 함축하지 않습니다. 모든 단계와 그 인수, 결과를 담은 전체 궤적은 내보낸 번들(oloproof export RUN_ID)과 워크벤치의 케이스 화면에 있습니다. 의미 있는 다음 조치는 애플리케이션에 있습니다. 절대 호출해서는 안 되는 도구를 루프가 호출하지 못하게 막으세요.

후보 변경을 만들고 비교하기

app.py에서 루프가 금지된 도구를 거부하게 하세요.

    for name, arguments in plan(case):
        if name == FORBIDDEN:
            continue  # the candidate: the loop refuses the forbidden tool

비교 정책 compare.yaml을 작성하세요.

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answers-not-worse
    metric: answer_correct
    kind: non_inferiority
    margin: 0.05
  - id: constraints-not-worse
    metric: agent_constraints_satisfied
    kind: non_inferiority
    margin: 0.05
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

후보 실행만 보면, no-deletion은 이제 PASS하고, agent_constraints_satisfied는 40 / 40을 읽으며, 두 하한 규칙이 여전히 INSUFFICIENT_EVIDENCE이므로 게이트는 종료 코드 3으로 차단합니다. 비교는 다음과 같습니다.

Comparison sha256:72836e90… of run_01M4FCFH0K2RRCN3ARAEN189X7 against run_01M4FCF6544JRDB16NJ1ZFPVRZ · 40 paired cases
answer_correct: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
agent_tool_lookup_order_called: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_no_tool_loop: +0.0 points [-17.7, +17.7] · 39 paired · 1 missing · 0 excluded
agent_tool_sequence: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_constraints_satisfied: +7.5 points [-7.8, +26.1] · 40 paired · 0 missing · 0 excluded
agent_steps_le_10: +0.0 points [-12.7, +12.7] · 40 paired · 0 missing · 0 excluded
steps_p95: +0 steps [+0, no bound] steps · p95 of per-case differences · 39 paired · 1 missing
tool_calls_p50: +0 calls [+0, +0] calls · p50 of per-case differences · 39 paired · 1 missing
72 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  constraints-not-worse  agent_constraints_satisfied  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 10 more paired cases would decide it, if the difference holds (50 in total at 8% discordance)
Gate: BLOCK (exit 3)

두 부분을 따로 읽으세요. 이 스위트에서 변경은 관측된 모든 삭제를 없앴고, 이는 단일 실행의 관측 개수 규칙이 이미 결론짓습니다. 후보가 일반적으로 기준선보다 나쁘지 않은지는 다른 질문이며, 짝지은 케이스 40개로는 아직 5포인트 마진 안에서 그것을 입증할 수 없습니다. 계획 줄은 대략 몇 개가 더 있으면 되는지 알려 줍니다. 바뀐 답은 없으므로, 수정은 작업 성공에 영향을 주지 않았습니다.

2부: 에이전트 팀

oloproof init --example triage_agents my-team
cd my-team

파일과 기록

app.py는 에이전트 셋을 하나의 루프에서 실행합니다. triage가 각 요청을 billing이나 tech에 넘기고, 각 전문 에이전트는 자신의 도구를 호출하며, billing이 발행할 수 없는 환불은 사람에게 넘겨집니다. 각 단계는 그것을 수행한 에이전트를 밝히고, 제어권의 각 이전은 handoff 단계입니다.

steps.append(AgentStep(index=1, kind="message", agent="triage", arguments={"request": request}))
steps.append(AgentStep(index=2, kind="handoff", agent="triage", to_agent="billing"))
steps.append(AgentStep(index=3, kind="tool_call", agent="billing", tool_name="lookup_order"))

궤적은 모든 단계의 에이전트를 밝히거나 하나도 밝히지 않아야 하며, 일부만 밝힌 궤적은 거부됩니다. 케이스는 따라야 할 경로를 선언합니다.

{"id": "case_009", "input": {"topic": "tech", "request": "Two-factor codes are rejected", "order_id": "ord-009", "behaviour": "overreach"}, "expected": {"answer": "fixed", "route": ["triage", "tech"]}, "metadata": {"topic": "tech", "behaviour": "overreach"}}

평가기와 정책

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: agent_route}
  - type: agent_tool_permissions
    permissions:
      triage: []
      billing: [lookup_order, issue_refund]
      tech: [search_kb]
  - {type: agent_max_handoffs, max_handoffs: 2}
slices: [route]
min_slice_support: 3
  • agent_route는 제어권을 가졌던 에이전트(반복은 합치고, 핸드오프의 수신자 포함)를 expected.route와 비교합니다. 성공 검사가 아니라 라우팅 검사입니다.
  • agent_tool_permissions는 모든 호출을 닫힌 맵에 대해 검사합니다. 맵에 없는 에이전트는 어떤 도구도 호출할 수 없습니다.
  • agent_max_handoffs는 제어권이 몇 번 넘어갔는지를 제한합니다.

release.yaml에는 answer-floor(min: 0.80), routing-floor(min: 0.70), 그리고 agent_tool_permissions에 대해 max_failures: 0인 관측 개수 규칙 no-overreach가 있습니다.

팀 실행하기

oloproof run
Gate: BLOCK (exit 1)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum      │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ no-overreach  │ agent_tool_permissions │ FAIL                  │ observed_failures_exceed_limit │
│ answer_correct         │ 96.7%    │ [82.7%, 100.0%] │ 29 / 30 observed · 0 missing · 0 excluded │
│ agent_route            │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 93.1%    │ [73.4%, 99.2%]  │ 27 / 29 observed · 1 missing · 0 excluded │
│ agent_handoffs_le_2    │ 86.7%    │ [69.2%, 96.3%]  │ 26 / 30 observed · 0 missing · 0 excluded │
oloproof inspect RUN_ID --failures
6 of 30 cases failed, errored or did not finish

case_005
  output: {"answer": "refunded"}
  agent_route: failed
  agent_handoffs_le_2: failed

case_009
  output: {"answer": "fixed"}
  agent_tool_permissions: failed
...
case_030
  output: {"answer": "unresolved"}
  answer_correct: failed
  agent_route: failed
  agent_tool_permissions: error: MissingFieldError: truncated_trajectory: the trace stops before whether an agent called a tool it was not given is settled
  agent_handoffs_le_2: failed
  • case_009와 case_020: tech가 환불을 발행했는데, 그 도구는 billing만 가집니다. 둘 다 올바르게 답했습니다. 작업은 성공, 권한은 위반입니다.
  • case_005와 다른 두 케이스는 처음에 잘못된 전문 에이전트로 갔다가 triage를 거쳐 돌아왔습니다. 답은 맞지만, 경로와 핸드오프 한계는 그렇지 않습니다.
  • case_030은 루프의 한계까지 billing과 tech 사이를 오갔습니다. 잘린 트레이스는 경로와 핸드오프 실패는 이미 증명하지만 권한은 결론지을 수 없으므로, 그 기준에서는 통과가 아니라 누락입니다.

출력의 어디에도 어느 에이전트의 탓인지는 나오지 않습니다. 경로 분기는 두 경로가 어디서 갈라지는지를 말해 줄 뿐입니다. 어떤 에이전트가 실패를 일으켰다는 것은 그 에이전트가 달리 행동했다면 무슨 일이 일어났을지에 대한 주장이며, 여기의 어떤 검사도 그런 주장을 하지 않습니다.

팀을 바꾸고 비교하기

권한 실패에 대한 의미 있는 다음 조치는 tech가 환불을 직접 발행하지 않고 billing에 넘기는 것입니다. app.py의 tech에서 다음과 같이 합니다.

        # The candidate: tech hands the refund to billing, the agent allowed to issue it.
        trace.hand_off("tech", "billing", "a goodwill refund")
        trace.call("billing", "issue_refund", order_id=str(case["order_id"]))

answer_correct에 대한 answers-not-worse와 agent_route에 대한 routing-not-worse를 담고, 둘 다 margin: 0.05인 non_inferiority로 둔 compare.yaml로 다음을 실행합니다.

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

후보 실행만 보면 다음과 같습니다.

Gate: BLOCK (exit 3)
│ answer-floor  │ answer_correct         │ PASS                  │ lower_bound_meets_minimum    │
│ routing-floor │ agent_route            │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold  │
│ no-overreach  │ agent_tool_permissions │ INSUFFICIENT_EVIDENCE │ missing_could_change_outcome │
│ agent_route            │ 80.0%    │ [61.4%, 92.3%]  │ 24 / 30 observed · 0 missing · 0 excluded │
│ agent_tool_permissions │ 100.0%   │ [82.7%, 100.0%] │ 29 / 29 observed · 1 missing · 0 excluded │

그리고 비교는 다음과 같습니다.

answer_correct: +0.0 points [-16.5, +16.5] · 30 paired · 0 missing · 0 excluded
agent_route: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
agent_tool_permissions: +6.9 points [-18.9, +33.5] · 29 paired · 1 missing · 0 excluded
agent_handoffs_le_2: -6.7 points [-28.5, +12.4] · 30 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  routing-not-worse  agent_route  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    no sample size would make this PASS: the difference itself (-6.7 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

여기서 얻을 세 가지는 다음과 같습니다.

  • 관측된 호출 중 권한을 어긴 것은 없지만, no-overreach는 이제 PASS가 아니라 INSUFFICIENT_EVIDENCE입니다. 잘린 case_030이 기록되지 않은 부분에 위반을 숨기고 있을 수 있기 때문입니다(missing_could_change_outcome). 규칙을 결론짓는 것은 권한 맵이 아니라 그 루프를 고치는 일입니다.
  • 수정으로 두 케이스의 경로가 triage > tech > billing으로 바뀌었는데, 그들의 expected.route는 이를 선언하지 않으므로 agent_route와 agent_handoffs_le_2가 떨어졌습니다. 그 경로가 이제 올바른지는 제품 결정입니다. 올바르다면 케이스의 expected.route를 갱신하세요. 라우팅 검사는 품질이 아니라 선언한 것에 대한 부합을 측정합니다.
  • 계획 줄은 케이스가 더 있으면 routing-not-worse가 PASS가 아니라 FAIL 쪽으로 움직인다고 말합니다. 비교는 현재 작성된 후보가 권한을 위해 라우팅을 희생한다고 알려 주고 있습니다.

문제 해결

증상원인해결
에이전트 평가기가 실행을 거부함시스템에 records: [agent_trajectory/v1]이 없음oloproof.yaml과 @system에 선언하세요
궤적이 거부됨일부 단계는 agent를 밝히고 다른 단계는 밝히지 않음모든 단계의 에이전트를 밝히거나 하나도 밝히지 마세요
한 기준에서 많은 케이스가 missing잘린 트레이스: 루프가 한계에 닿음한계를 높이거나 루프를 고치세요. 누락 케이스는 통과하지 않고 구간을 넓힙니다
agent_tool_sequence의 분모가 작음expected.tools가 없는 케이스중요한 곳에 순서를 선언하세요. []은 "도구를 기대하지 않음"을 뜻합니다
제약 지표가 절대 실패하지 않음애플리케이션이 그 검사를 기록하지 않음환경이 관찰하는 곳에서 AgentConstraintCheck를 기록하세요
같은 버전의 실행 사이에 결과가 다름도구가 공유 상태를 읽거나 씀애플리케이션에서 각 케이스 전에 그 상태를 초기화하세요. Oloproof는 하지 않습니다
팀에 추가한 에이전트가 곧바로 권한에서 실패함권한 맵이 닫혀 있음새 에이전트가 호출할 수 있는 것을 선언하세요

제한 사항

  • Oloproof는 에이전트를 구동하거나, 샌드박스에 넣거나, 초기화하지 않습니다. 도구의 부수 효과, 세션, 상태와 그 초기화는 애플리케이션의 몫입니다.
  • 모든 검사는 기록된 궤적을 읽습니다. 애플리케이션이 기록하지 않은 것은 측정할 수 없으며, 잘린 트레이스는 그 앞부분이 질문을 결론짓지 못하는 곳마다 누락으로 셉니다.
  • 궤적 검사는 결정론적 규칙입니다. LLM이 심사하는 궤적 품질 검사는 없습니다.
  • 실패를 단계나 에이전트에 귀속시키는 출력은 없습니다. 기록된 체크포인트에서 단계 하나를 빼고 케이스를 다시 실행해 그 단계가 필요한지 불필요한지 레이블을 붙이는 에이전트 재생은, 체크포인트로부터의 재생을 구현한 시스템에 한해 Python SDK에만(replay_case) 있습니다. 그에 대한 CLI 명령은 없으며, 위의 어떤 것도 그것을 쓰지 않습니다.
  • 여러 턴의 대화는 다른 영역입니다(SDK 전용). 지금 작동하는 것을 보세요.
  • 예제의 plan과 behaviour 필드는 실행을 재현할 수 있도록 모델의 결정을 대신합니다. 루프 안의 라이브 모델은 제공자를 호출하고, 자격 증명이 필요하며, 케이스마다 비용이 듭니다.