跳至主要內容

指南

SDK 參考

oloproof 和 oloproof.evaluators 包導出的每個名稱,附帶其簽名、它是同步還是非同步的,以及它返回什麼。如需引導式介紹,請先閱讀 Python API。

只有這兩個包是公開接口。從 oloproof_core 導入的任何東西都是引擎內部實作,可能在不另行通知的情況下改變。下面的每個函式都在本機針對專案的儲存執行;除非你傳入的某個評估器呼叫了模型提供者,否則它們都不會把資料發送到任何地方。

執行評估

evaluate 與 aevaluate

def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
參數類型它是什麼
system一個 @system 函式、一個 @rag_system 類或實例,或一個可呼叫物件被測系統。
dataset路徑一個 JSONL 套件。見套件。
evaluators清單來自 oloproof.evaluators 的實例,或 @evaluator 函式。
policyrelease.yaml 的路徑、一個 ReleasePolicy,或 None發布政策。None 不執行閘門:result.gate 為 None,什麼都不決策。
concurrencyConcurrencyConfig 或像 {"system": 8, "judge": 4} 這樣的映射同時進行的呼叫數。
slices字串清單探索性切片,與 oloproof.yaml 中相同。
min_slice_support整數符合條件的案例少於這個數時,切片沒有區間。預設 30。
replicates整數每個案例量測這麼多次。預設 1。

evaluate 是同步的。在沒有執行中的事件循環時呼叫,它使用 asyncio.run;在執行中的循環內部呼叫時(筆記本、非同步測試),它在一個單獨的線程上執行評估並阻塞直到完成,所以在兩種場合都可以安全使用。aevaluate 是協程;在非同步代碼中 await 它。

其餘的關鍵字參數(metrics、store、predictive、event_sink、retry_policy、traffic_draw_id)接受來自 oloproof_core 的引擎類型,不屬於穩定接口。

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator

@system(name="support-bot", version="1")
def answer(case):
    current_case().usage(input_tokens=12, output_tokens=3)
    return {"label": "refund" if "refund" in case["question"].lower() else "other"}

@evaluator(criterion="short_label")
def short_label(case):
    return len(case.output["label"]) <= 6

result = evaluate(
    system=answer,
    dataset="cases.jsonl",
    evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)

在一個兩案例的套件上不帶政策執行,印出出:

correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0

EvaluationResult

成員類型它是什麼
run執行記錄已儲存的執行,帶有它的 id、狀態和完整性。
suite、system、evaluators版本記錄這次執行所量測的確切版本。
metrics指標結果的元組每個判據和宣告的指標一個:metric、estimate、interval、n_total、n_eligible、n_observed、n_missing、exclusions、method。
gate閘門結果或 None有政策時:release_action、exit_code、decisions 和 reasons。
decisions元組閘門的決策,沒有政策時為空。
cases()清單每個案例及其執行和評判結果。
failures()清單沒有完成、出錯或至少在一個評估器上失敗的案例。
print(stderr=False)無oloproof run 印出的終端機報告。
to_bundle(path)路徑寫出一個可移植的包,與 oloproof export 相同。

狀態、原因代碼和計數的含義見結果與執行。

evaluate_comparison 與 aevaluate_comparison

def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult

以同一個帶種子的順序在同一個套件上執行兩個系統,並對政策中的比較規則作出決策。policy 是必需的,並且必須至少包含一條 superiority、non_inferiority 或 equivalence 規則,否則呼叫會拋出設定錯誤。額外的關鍵字參數:concurrency、replicates,以及引擎類型的 metrics、store 和 retry_policy。同步和非同步的行為與 evaluate 相同。

ComparisonEvaluationResult 包含 candidate 和 baseline(各為一個 EvaluationResult)以及 comparison,後者帶有配對差異及其決策。見比較兩個版本。

宣告系統

@system

def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())

可以直接使用(@system),也可以帶參數使用(@system(name=..., version=...)),或者對一個物件呼叫(system(model.answer, version="v2"))。函式接收的是案例的 input 物件,而不是整個案例,並返回評估器讀取的輸出。它可以是 def 或 async def;同步函式在工作線程上執行。

參數預設值它是什麼
name函式的名稱版本身份的一部分。
version無對於綁定方法或可呼叫物件是必需的,因為它們的行為取決於 Oloproof 看不到的狀態。
config空與版本一同記錄的設置。
timeout_s120每次呼叫的時限。超時的呼叫會被記錄為超時的執行。
records空系統記錄的產物類型,例如 retrieval/v1。如果某個評估器需要一個系統沒有宣告的類型,它會在執行開始之前被拒絕。

函式自身模組的原始碼會計入版本摘要,所以編輯它會使快取的執行失效。還有哪些會、哪些不會,見設定參考。

current_case

def current_case() -> CaseRecorder

只有在 Oloproof 呼叫你的系統期間才可用;在其他任何地方它都會拋出 RuntimeError。記錄器的方法:

方法記錄
usage(*, input_tokens=None, output_tokens=None, cost_usd=None)一次模型呼叫的 token 數和成本。省略的值保持未記錄,而不是零。
artifact(kind, data)任意 JSON 值或 Pydantic 模型,歸入 trace 或 conversation/v1 這樣的類型下。
retrieval(retrieval)檢索器返回的排序後的候選(retrieval/v1)。
context(context)為生成組裝的上下文(context/v1)。
citations(ids)回答引用的 id,形式為 doc_id 或 doc_id#chunk_id(citations/v1)。
agent_trajectory(trajectory)代理的步驟、工具呼叫和結果,以及檢查點(agent_trajectory/v1)。

每個方法都返回一個 ArtifactRef(usage 除外,它什麼都不返回)。為了建置這些記錄,導出了以下有類型的載荷:Retrieval、Passage、Context、ContextItem、DroppedItem、Citations、StageTimings、AgentTrajectory、AgentStep、AgentCheckpoint、AgentConstraintCheck,以及類型名 CONVERSATION(conversation/v1)。

@rag_system

def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")

一個類裝飾器。該類提供 retrieve(input, depth) 和 generate(input, context),設置了 token_budget 時還要提供 count_tokens(passage)。context 是經過 top_k 和預算篩選後留下的 Passage 物件清單,按排名順序排列。Oloproof 自己記錄 retrieval/v1、context/v1、citations/v1 和 stage_timings/v1,並分別快取每個階段。citations_path 指明保存回答所引用 id 的輸出欄位。見 RAG。

評估器

所有類都在 oloproof.evaluators 中。每個類的 criterion 指明它產生的指標。每個評估器讀取哪些產物,以及它在 YAML 中的對應寫法,見設定參考中的評估器表。

類簽名
ExactMatch(*, criterion, field=None, expected_field=None, strip=True, casefold=False)
Contains(*, criterion, field=None, expected_field=None)
Regex(*, criterion, pattern, field=None, pass_if="match")
JsonSchema(*, criterion, schema, field=None)
RubricJudge(*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0)
Groundedness(*, provider, model, criterion="groundedness", **options)
CitationSupport(*, provider, model, criterion="citation_support", **options)
CitationValidity(*, criterion="citations_valid", require_citations=False)
HitRate、Recall(k=None, *, criterion=None, relevance_unit="doc"),k 預設為 5
MRR、NDCG(k=None, *, criterion=None, relevance_unit="doc"),k 預設為 10
AgentMaxSteps(max_steps, *, criterion=None)
AgentToolCalled(tool_name, *, min_calls=1, criterion=None)
AgentNoToolLoop(*, max_repeats=2, criterion="agent_no_tool_loop")
AgentToolSequence(*, ordered=True, criterion="agent_tool_sequence")
AgentNoUndeclaredTool(*, criterion="agent_no_undeclared_tool")
AgentConstraintsSatisfied(constraints=(), *, criterion="agent_constraints_satisfied")
AgentRoute(*, criterion="agent_route")
AgentToolPermissions(permissions, *, criterion="agent_tool_permissions")
AgentMaxHandoffs(max_handoffs, *, criterion=None)
ConversationCompleted(*, criterion="conversation_completed")
ConversationJudge與 RubricJudge 相同
PredictiveCorrect、PredictiveRecall、PredictivePrecision(*, criterion, positive=True, field="label", expected_field="label")
AbsoluteError(*, criterion, target_range, field="label", expected_field="label")
Brier、PredictiveRanking(*, criterion, positive=True, field="score", expected_field="label")
LogLoss(*, clip, criterion, positive=True, field="score", expected_field="label")
CustomEvaluator(func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

provider 是 "anthropic"、"openai" 或 "openai_compatible"。評審從 api_key_env 指定的環境變數中讀取密鑰(預設為 ANTHROPIC_API_KEY 或 OPENAI_API_KEY),並由該提供者計費。Groundedness 和 CitationSupport 通過 **options 接受 RubricJudge 的其餘設置。機率評審、模型分類器和級聯沒有 SDK 類;它們只存在於 YAML 中。

ConversationCompleted 和 ConversationJudge 讀取你的系統記錄的 conversation/v1 產物。Oloproof 不驅動對話:你的應用執行每一輪並記錄對話記錄。見代理。

@evaluator

def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

把一個單參數(即案例)的函式包裝成一個 CustomEvaluator。案例具有 output、expected 和 scenario,而 artifacts(name) 返回某個已記錄類型的載荷。函式可以是 def 或 async def。二值評估器返回 True 或 False;分數評估器宣告 value_type="score" 和 score_range=(low, high),並返回一個數字。函式拋出的異常會把該案例在這個判據上記為缺失,而絕不會記為失敗。

reads 必須列出函式讀取的每個欄位(input、output、expected、metadata、metadata.<key> 或 artifacts.<name>),因為快取的評判結果正是以這些欄位為鍵的。只有在 cacheable=True 時,評判結果才會在多次執行之間被複用。定義模組的原始碼會計入版本,所以編輯它會使這些評判結果失效。YAML 無法指定自定義評估器。

診斷

diagnose 與 adiagnose

def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResult

在一種干預下重新執行已儲存執行中失敗的案例:"gold-context"、"top-k"(配合 top_k)或 "reranker"(配合 reranker)。system 和 evaluators 必須是該執行所用的版本;不同的版本會在任何執行之前被拒絕。使用 control=True 時,會在干預旁邊執行一個全新的對照樣本,從而把變更與執行之間的波動區分開。diagnose 是同步形式,在執行中的循環內部的行為與 evaluate 相同。工作流程見 RAG。

InterventionResult 保存父執行 id、干預、它是否受支援、干預執行和對照執行、每個案例的結果,以及在生成了報告時由 CaseDiagnosis 條目組成的 DiagnosisReport。

代理重放

def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]

重放是你的系統做的事,而不是 Oloproof 模擬的事。系統只有通過實作 async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory,返回代理從檢查點開始所做的事,才能支援重放;Oloproof 把記錄下來的前綴拼接上去並進行比較。你的應用負責其狀態、會話和工具副作用,包括在重放之前重置它們。supports_replay 報告一個系統是否宣告瞭該方法。

replay_case 是一個協程:await 它,或通過 asyncio.run 呼叫它。它在重放之前執行一個對照(使用 ReplayChange(kind="resume") 的同一個檢查點),當對照無法復現記錄時,不會嘗試重放。change 是 ReplayChange(kind="drop_step", step_index=...) 或 ReplayChange(kind="resume")。checkpoint_for 選取嚴格位於某一步之前的最新記錄檢查點,或者 None,在後一種情況下,結果會以 no_checkpoint_recorded 被丟棄,而不觸及系統。

ReplayOutcome 帶有 scenario_id、change、重放和對照的終止狀態,以及當該案例沒有產生證據時帶原因的 discarded。被丟棄的案例不是失敗的案例。label_case 把一個結果轉換為帶有 FailureLabel 及其 LabelReason 的 CaseDiagnosis;unnecessary_steps 列出去掉後結果不變的步驟。見代理。

人工標籤與評估器信任

def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabel

儲存一個人對已儲存執行中某個案例給出的通過/失敗結論。標籤供 oloproof evaluators validate 使用,它量測評審與標籤的一致性(AgreementResult)並記錄其狀態(RegistryEntry)。其餘的 measurement_sample_* 參數把標籤綁定到一個量測樣本;CLI 的 oloproof labels export 和 oloproof labels import 會為你填寫它們。見評審。

在代碼中編寫政策

ReleasePolicy、IntervalThresholdRule、ObservedCountRule 和 DecisionRule(前兩者的聯合)無需檔案即可建置政策。它們的欄位就是設定參考中 release.yaml 的欄位;區間規則使用 direction(min 或 max)和 threshold,而不是 min: 或 max:。比較規則沒有導出的類;把它們寫在 release.yaml 中並傳入其路徑。

from oloproof import IntervalThresholdRule, ReleasePolicy

policy = ReleasePolicy(
    rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)

其他導出

名稱它是什麼
ConcurrencyConfigsystem 和 judge 的限制,與 oloproof.yaml 中相同。
TransientError在系統中帶 retryable=True 拋出它,讓呼叫以退避方式重試。
ArtifactRef記錄產物時返回的引用:它的類型和摘要。
__version__已安裝的軟體包版本。