指南
SDK 參考
oloproof 和 oloproof.evaluators 包導出的每個名稱,附帶其簽名、它是同步還是非同步的,以及它返回什麼。如需引導式介紹,請先閱讀 Python API。
只有這兩個包是公開接口。從 oloproof_core 導入的任何東西都是引擎內部實作,可能在不另行通知的情況下改變。下面的每個函式都在本機針對專案的儲存執行;除非你傳入的某個評估器呼叫了模型提供者,否則它們都不會把資料發送到任何地方。
執行評估
evaluate 與 aevaluate
def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult| 參數 | 類型 | 它是什麼 |
|---|---|---|
| system | 一個 @system 函式、一個 @rag_system 類或實例,或一個可呼叫物件 | 被測系統。 |
| dataset | 路徑 | 一個 JSONL 套件。見套件。 |
| evaluators | 清單 | 來自 oloproof.evaluators 的實例,或 @evaluator 函式。 |
| policy | release.yaml 的路徑、一個 ReleasePolicy,或 None | 發布政策。None 不執行閘門:result.gate 為 None,什麼都不決策。 |
| concurrency | ConcurrencyConfig 或像 {"system": 8, "judge": 4} 這樣的映射 | 同時進行的呼叫數。 |
| slices | 字串清單 | 探索性切片,與 oloproof.yaml 中相同。 |
| min_slice_support | 整數 | 符合條件的案例少於這個數時,切片沒有區間。預設 30。 |
| replicates | 整數 | 每個案例量測這麼多次。預設 1。 |
evaluate 是同步的。在沒有執行中的事件循環時呼叫,它使用 asyncio.run;在執行中的循環內部呼叫時(筆記本、非同步測試),它在一個單獨的線程上執行評估並阻塞直到完成,所以在兩種場合都可以安全使用。aevaluate 是協程;在非同步代碼中 await 它。
其餘的關鍵字參數(metrics、store、predictive、event_sink、retry_policy、traffic_draw_id)接受來自 oloproof_core 的引擎類型,不屬於穩定接口。
from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator
@system(name="support-bot", version="1")
def answer(case):
current_case().usage(input_tokens=12, output_tokens=3)
return {"label": "refund" if "refund" in case["question"].lower() else "other"}
@evaluator(criterion="short_label")
def short_label(case):
return len(case.output["label"]) <= 6
result = evaluate(
system=answer,
dataset="cases.jsonl",
evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)在一個兩案例的套件上不帶政策執行,印出出:
correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0EvaluationResult
| 成員 | 類型 | 它是什麼 |
|---|---|---|
| run | 執行記錄 | 已儲存的執行,帶有它的 id、狀態和完整性。 |
| suite、system、evaluators | 版本記錄 | 這次執行所量測的確切版本。 |
| metrics | 指標結果的元組 | 每個判據和宣告的指標一個:metric、estimate、interval、n_total、n_eligible、n_observed、n_missing、exclusions、method。 |
| gate | 閘門結果或 None | 有政策時:release_action、exit_code、decisions 和 reasons。 |
| decisions | 元組 | 閘門的決策,沒有政策時為空。 |
| cases() | 清單 | 每個案例及其執行和評判結果。 |
| failures() | 清單 | 沒有完成、出錯或至少在一個評估器上失敗的案例。 |
| print(stderr=False) | 無 | oloproof run 印出的終端機報告。 |
| to_bundle(path) | 路徑 | 寫出一個可移植的包,與 oloproof export 相同。 |
狀態、原因代碼和計數的含義見結果與執行。
evaluate_comparison 與 aevaluate_comparison
def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult以同一個帶種子的順序在同一個套件上執行兩個系統,並對政策中的比較規則作出決策。policy 是必需的,並且必須至少包含一條 superiority、non_inferiority 或 equivalence 規則,否則呼叫會拋出設定錯誤。額外的關鍵字參數:concurrency、replicates,以及引擎類型的 metrics、store 和 retry_policy。同步和非同步的行為與 evaluate 相同。
ComparisonEvaluationResult 包含 candidate 和 baseline(各為一個 EvaluationResult)以及 comparison,後者帶有配對差異及其決策。見比較兩個版本。
宣告系統
@system
def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())可以直接使用(@system),也可以帶參數使用(@system(name=..., version=...)),或者對一個物件呼叫(system(model.answer, version="v2"))。函式接收的是案例的 input 物件,而不是整個案例,並返回評估器讀取的輸出。它可以是 def 或 async def;同步函式在工作線程上執行。
| 參數 | 預設值 | 它是什麼 |
|---|---|---|
| name | 函式的名稱 | 版本身份的一部分。 |
| version | 無 | 對於綁定方法或可呼叫物件是必需的,因為它們的行為取決於 Oloproof 看不到的狀態。 |
| config | 空 | 與版本一同記錄的設置。 |
| timeout_s | 120 | 每次呼叫的時限。超時的呼叫會被記錄為超時的執行。 |
| records | 空 | 系統記錄的產物類型,例如 retrieval/v1。如果某個評估器需要一個系統沒有宣告的類型,它會在執行開始之前被拒絕。 |
函式自身模組的原始碼會計入版本摘要,所以編輯它會使快取的執行失效。還有哪些會、哪些不會,見設定參考。
current_case
def current_case() -> CaseRecorder只有在 Oloproof 呼叫你的系統期間才可用;在其他任何地方它都會拋出 RuntimeError。記錄器的方法:
| 方法 | 記錄 |
|---|---|
| usage(*, input_tokens=None, output_tokens=None, cost_usd=None) | 一次模型呼叫的 token 數和成本。省略的值保持未記錄,而不是零。 |
| artifact(kind, data) | 任意 JSON 值或 Pydantic 模型,歸入 trace 或 conversation/v1 這樣的類型下。 |
| retrieval(retrieval) | 檢索器返回的排序後的候選(retrieval/v1)。 |
| context(context) | 為生成組裝的上下文(context/v1)。 |
| citations(ids) | 回答引用的 id,形式為 doc_id 或 doc_id#chunk_id(citations/v1)。 |
| agent_trajectory(trajectory) | 代理的步驟、工具呼叫和結果,以及檢查點(agent_trajectory/v1)。 |
每個方法都返回一個 ArtifactRef(usage 除外,它什麼都不返回)。為了建置這些記錄,導出了以下有類型的載荷:Retrieval、Passage、Context、ContextItem、DroppedItem、Citations、StageTimings、AgentTrajectory、AgentStep、AgentCheckpoint、AgentConstraintCheck,以及類型名 CONVERSATION(conversation/v1)。
@rag_system
def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")一個類裝飾器。該類提供 retrieve(input, depth) 和 generate(input, context),設置了 token_budget 時還要提供 count_tokens(passage)。context 是經過 top_k 和預算篩選後留下的 Passage 物件清單,按排名順序排列。Oloproof 自己記錄 retrieval/v1、context/v1、citations/v1 和 stage_timings/v1,並分別快取每個階段。citations_path 指明保存回答所引用 id 的輸出欄位。見 RAG。
評估器
所有類都在 oloproof.evaluators 中。每個類的 criterion 指明它產生的指標。每個評估器讀取哪些產物,以及它在 YAML 中的對應寫法,見設定參考中的評估器表。
| 類 | 簽名 |
|---|---|
| ExactMatch | (*, criterion, field=None, expected_field=None, strip=True, casefold=False) |
| Contains | (*, criterion, field=None, expected_field=None) |
| Regex | (*, criterion, pattern, field=None, pass_if="match") |
| JsonSchema | (*, criterion, schema, field=None) |
| RubricJudge | (*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0) |
| Groundedness | (*, provider, model, criterion="groundedness", **options) |
| CitationSupport | (*, provider, model, criterion="citation_support", **options) |
| CitationValidity | (*, criterion="citations_valid", require_citations=False) |
| HitRate、Recall | (k=None, *, criterion=None, relevance_unit="doc"),k 預設為 5 |
| MRR、NDCG | (k=None, *, criterion=None, relevance_unit="doc"),k 預設為 10 |
| AgentMaxSteps | (max_steps, *, criterion=None) |
| AgentToolCalled | (tool_name, *, min_calls=1, criterion=None) |
| AgentNoToolLoop | (*, max_repeats=2, criterion="agent_no_tool_loop") |
| AgentToolSequence | (*, ordered=True, criterion="agent_tool_sequence") |
| AgentNoUndeclaredTool | (*, criterion="agent_no_undeclared_tool") |
| AgentConstraintsSatisfied | (constraints=(), *, criterion="agent_constraints_satisfied") |
| AgentRoute | (*, criterion="agent_route") |
| AgentToolPermissions | (permissions, *, criterion="agent_tool_permissions") |
| AgentMaxHandoffs | (max_handoffs, *, criterion=None) |
| ConversationCompleted | (*, criterion="conversation_completed") |
| ConversationJudge | 與 RubricJudge 相同 |
| PredictiveCorrect、PredictiveRecall、PredictivePrecision | (*, criterion, positive=True, field="label", expected_field="label") |
| AbsoluteError | (*, criterion, target_range, field="label", expected_field="label") |
| Brier、PredictiveRanking | (*, criterion, positive=True, field="score", expected_field="label") |
| LogLoss | (*, clip, criterion, positive=True, field="score", expected_field="label") |
| CustomEvaluator | (func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None) |
provider 是 "anthropic"、"openai" 或 "openai_compatible"。評審從 api_key_env 指定的環境變數中讀取密鑰(預設為 ANTHROPIC_API_KEY 或 OPENAI_API_KEY),並由該提供者計費。Groundedness 和 CitationSupport 通過 **options 接受 RubricJudge 的其餘設置。機率評審、模型分類器和級聯沒有 SDK 類;它們只存在於 YAML 中。
ConversationCompleted 和 ConversationJudge 讀取你的系統記錄的 conversation/v1 產物。Oloproof 不驅動對話:你的應用執行每一輪並記錄對話記錄。見代理。
@evaluator
def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)把一個單參數(即案例)的函式包裝成一個 CustomEvaluator。案例具有 output、expected 和 scenario,而 artifacts(name) 返回某個已記錄類型的載荷。函式可以是 def 或 async def。二值評估器返回 True 或 False;分數評估器宣告 value_type="score" 和 score_range=(low, high),並返回一個數字。函式拋出的異常會把該案例在這個判據上記為缺失,而絕不會記為失敗。
reads 必須列出函式讀取的每個欄位(input、output、expected、metadata、metadata.<key> 或 artifacts.<name>),因為快取的評判結果正是以這些欄位為鍵的。只有在 cacheable=True 時,評判結果才會在多次執行之間被複用。定義模組的原始碼會計入版本,所以編輯它會使這些評判結果失效。YAML 無法指定自定義評估器。
診斷
diagnose 與 adiagnose
def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResult在一種干預下重新執行已儲存執行中失敗的案例:"gold-context"、"top-k"(配合 top_k)或 "reranker"(配合 reranker)。system 和 evaluators 必須是該執行所用的版本;不同的版本會在任何執行之前被拒絕。使用 control=True 時,會在干預旁邊執行一個全新的對照樣本,從而把變更與執行之間的波動區分開。diagnose 是同步形式,在執行中的循環內部的行為與 evaluate 相同。工作流程見 RAG。
InterventionResult 保存父執行 id、干預、它是否受支援、干預執行和對照執行、每個案例的結果,以及在生成了報告時由 CaseDiagnosis 條目組成的 DiagnosisReport。
代理重放
def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]重放是你的系統做的事,而不是 Oloproof 模擬的事。系統只有通過實作 async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory,返回代理從檢查點開始所做的事,才能支援重放;Oloproof 把記錄下來的前綴拼接上去並進行比較。你的應用負責其狀態、會話和工具副作用,包括在重放之前重置它們。supports_replay 報告一個系統是否宣告瞭該方法。
replay_case 是一個協程:await 它,或通過 asyncio.run 呼叫它。它在重放之前執行一個對照(使用 ReplayChange(kind="resume") 的同一個檢查點),當對照無法復現記錄時,不會嘗試重放。change 是 ReplayChange(kind="drop_step", step_index=...) 或 ReplayChange(kind="resume")。checkpoint_for 選取嚴格位於某一步之前的最新記錄檢查點,或者 None,在後一種情況下,結果會以 no_checkpoint_recorded 被丟棄,而不觸及系統。
ReplayOutcome 帶有 scenario_id、change、重放和對照的終止狀態,以及當該案例沒有產生證據時帶原因的 discarded。被丟棄的案例不是失敗的案例。label_case 把一個結果轉換為帶有 FailureLabel 及其 LabelReason 的 CaseDiagnosis;unnecessary_steps 列出去掉後結果不變的步驟。見代理。
人工標籤與評估器信任
def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabel儲存一個人對已儲存執行中某個案例給出的通過/失敗結論。標籤供 oloproof evaluators validate 使用,它量測評審與標籤的一致性(AgreementResult)並記錄其狀態(RegistryEntry)。其餘的 measurement_sample_* 參數把標籤綁定到一個量測樣本;CLI 的 oloproof labels export 和 oloproof labels import 會為你填寫它們。見評審。
在代碼中編寫政策
ReleasePolicy、IntervalThresholdRule、ObservedCountRule 和 DecisionRule(前兩者的聯合)無需檔案即可建置政策。它們的欄位就是設定參考中 release.yaml 的欄位;區間規則使用 direction(min 或 max)和 threshold,而不是 min: 或 max:。比較規則沒有導出的類;把它們寫在 release.yaml 中並傳入其路徑。
from oloproof import IntervalThresholdRule, ReleasePolicy
policy = ReleasePolicy(
rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)其他導出
| 名稱 | 它是什麼 |
|---|---|
| ConcurrencyConfig | system 和 judge 的限制,與 oloproof.yaml 中相同。 |
| TransientError | 在系統中帶 retryable=True 拋出它,讓呼叫以退避方式重試。 |
| ArtifactRef | 記錄產物時返回的引用:它的類型和摘要。 |
| __version__ | 已安裝的軟體包版本。 |