跳到主要内容

指南

SDK 参考

oloproof 和 oloproof.evaluators 包导出的每个名称,附带其签名、它是同步还是异步的,以及它返回什么。如需引导式介绍,请先阅读 Python API。

只有这两个包是公开接口。从 oloproof_core 导入的任何东西都是引擎内部实现,可能在不另行通知的情况下改变。下面的每个函数都在本地针对项目的存储运行;除非你传入的某个评估器调用了模型提供方,否则它们都不会把数据发送到任何地方。

运行评估

evaluate 与 aevaluate

def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
参数类型它是什么
system一个 @system 函数、一个 @rag_system 类或实例,或一个可调用对象被测系统。
dataset路径一个 JSONL 测试套件。见测试套件。
evaluators列表来自 oloproof.evaluators 的实例,或 @evaluator 函数。
policyrelease.yaml 的路径、一个 ReleasePolicy,或 None发布策略。None 不运行门禁:result.gate 为 None,什么都不决策。
concurrencyConcurrencyConfig 或像 {"system": 8, "judge": 4} 这样的映射同时进行的调用数。
slices字符串列表探索性切片,与 oloproof.yaml 中相同。
min_slice_support整数符合条件的用例少于这个数时,切片没有区间。默认 30。
replicates整数每个用例度量这么多次。默认 1。

evaluate 是同步的。在没有运行中的事件循环时调用,它使用 asyncio.run;在运行中的循环内部调用时(笔记本、异步测试),它在一个单独的线程上运行评估并阻塞直到完成,所以在两种场合都可以安全使用。aevaluate 是协程;在异步代码中 await 它。

其余的关键字参数(metrics、store、predictive、event_sink、retry_policy、traffic_draw_id)接受来自 oloproof_core 的引擎类型,不属于稳定接口。

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator

@system(name="support-bot", version="1")
def answer(case):
    current_case().usage(input_tokens=12, output_tokens=3)
    return {"label": "refund" if "refund" in case["question"].lower() else "other"}

@evaluator(criterion="short_label")
def short_label(case):
    return len(case.output["label"]) <= 6

result = evaluate(
    system=answer,
    dataset="cases.jsonl",
    evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)

在一个两用例的测试套件上不带策略运行,打印出:

correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0

EvaluationResult

成员类型它是什么
run运行记录已存储的运行,带有它的 id、状态和完整性。
suite、system、evaluators版本记录这次运行所度量的确切版本。
metrics指标结果的元组每个判据和声明的指标一个:metric、estimate、interval、n_total、n_eligible、n_observed、n_missing、exclusions、method。
gate门禁结果或 None有策略时:release_action、exit_code、decisions 和 reasons。
decisions元组门禁的决策,没有策略时为空。
cases()列表每个用例及其执行和评判结果。
failures()列表没有完成、出错或至少在一个评估器上失败的用例。
print(stderr=False)无oloproof run 打印的终端报告。
to_bundle(path)路径写出一个可移植的包,与 oloproof export 相同。

状态、原因代码和计数的含义见结果与执行。

evaluate_comparison 与 aevaluate_comparison

def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult

以同一个带种子的顺序在同一个测试套件上运行两个系统,并对策略中的比较规则作出决策。policy 是必需的,并且必须至少包含一条 superiority、non_inferiority 或 equivalence 规则,否则调用会抛出配置错误。额外的关键字参数:concurrency、replicates,以及引擎类型的 metrics、store 和 retry_policy。同步和异步的行为与 evaluate 相同。

ComparisonEvaluationResult 包含 candidate 和 baseline(各为一个 EvaluationResult)以及 comparison,后者带有配对差异及其决策。见比较两个版本。

声明系统

@system

def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())

可以直接使用(@system),也可以带参数使用(@system(name=..., version=...)),或者对一个对象调用(system(model.answer, version="v2"))。函数接收的是用例的 input 对象,而不是整个用例,并返回评估器读取的输出。它可以是 def 或 async def;同步函数在工作线程上运行。

参数默认值它是什么
name函数的名称版本身份的一部分。
version无对于绑定方法或可调用对象是必需的,因为它们的行为取决于 Oloproof 看不到的状态。
config空与版本一同记录的设置。
timeout_s120每次调用的时限。超时的调用会被记录为超时的执行。
records空系统记录的产物类型,例如 retrieval/v1。如果某个评估器需要一个系统没有声明的类型,它会在运行开始之前被拒绝。

函数自身模块的源代码会计入版本摘要,所以编辑它会使缓存的执行失效。还有哪些会、哪些不会,见配置参考。

current_case

def current_case() -> CaseRecorder

只有在 Oloproof 调用你的系统期间才可用;在其他任何地方它都会抛出 RuntimeError。记录器的方法:

方法记录
usage(*, input_tokens=None, output_tokens=None, cost_usd=None)一次模型调用的 token 数和成本。省略的值保持未记录,而不是零。
artifact(kind, data)任意 JSON 值或 Pydantic 模型,归入 trace 或 conversation/v1 这样的类型下。
retrieval(retrieval)检索器返回的排序后的候选(retrieval/v1)。
context(context)为生成组装的上下文(context/v1)。
citations(ids)回答引用的 id,形式为 doc_id 或 doc_id#chunk_id(citations/v1)。
agent_trajectory(trajectory)智能体的步骤、工具调用和结果,以及检查点(agent_trajectory/v1)。

每个方法都返回一个 ArtifactRef(usage 除外,它什么都不返回)。为了构建这些记录,导出了以下有类型的载荷:Retrieval、Passage、Context、ContextItem、DroppedItem、Citations、StageTimings、AgentTrajectory、AgentStep、AgentCheckpoint、AgentConstraintCheck,以及类型名 CONVERSATION(conversation/v1)。

@rag_system

def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")

一个类装饰器。该类提供 retrieve(input, depth) 和 generate(input, context),设置了 token_budget 时还要提供 count_tokens(passage)。context 是经过 top_k 和预算筛选后留下的 Passage 对象列表,按排名顺序排列。Oloproof 自己记录 retrieval/v1、context/v1、citations/v1 和 stage_timings/v1,并分别缓存每个阶段。citations_path 指明保存回答所引用 id 的输出字段。见 RAG。

评估器

所有类都在 oloproof.evaluators 中。每个类的 criterion 指明它产生的指标。每个评估器读取哪些产物,以及它在 YAML 中的对应写法,见配置参考中的评估器表。

类签名
ExactMatch(*, criterion, field=None, expected_field=None, strip=True, casefold=False)
Contains(*, criterion, field=None, expected_field=None)
Regex(*, criterion, pattern, field=None, pass_if="match")
JsonSchema(*, criterion, schema, field=None)
RubricJudge(*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0)
Groundedness(*, provider, model, criterion="groundedness", **options)
CitationSupport(*, provider, model, criterion="citation_support", **options)
CitationValidity(*, criterion="citations_valid", require_citations=False)
HitRate、Recall(k=None, *, criterion=None, relevance_unit="doc"),k 默认为 5
MRR、NDCG(k=None, *, criterion=None, relevance_unit="doc"),k 默认为 10
AgentMaxSteps(max_steps, *, criterion=None)
AgentToolCalled(tool_name, *, min_calls=1, criterion=None)
AgentNoToolLoop(*, max_repeats=2, criterion="agent_no_tool_loop")
AgentToolSequence(*, ordered=True, criterion="agent_tool_sequence")
AgentNoUndeclaredTool(*, criterion="agent_no_undeclared_tool")
AgentConstraintsSatisfied(constraints=(), *, criterion="agent_constraints_satisfied")
AgentRoute(*, criterion="agent_route")
AgentToolPermissions(permissions, *, criterion="agent_tool_permissions")
AgentMaxHandoffs(max_handoffs, *, criterion=None)
ConversationCompleted(*, criterion="conversation_completed")
ConversationJudge与 RubricJudge 相同
PredictiveCorrect、PredictiveRecall、PredictivePrecision(*, criterion, positive=True, field="label", expected_field="label")
AbsoluteError(*, criterion, target_range, field="label", expected_field="label")
Brier、PredictiveRanking(*, criterion, positive=True, field="score", expected_field="label")
LogLoss(*, clip, criterion, positive=True, field="score", expected_field="label")
CustomEvaluator(func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

provider 是 "anthropic"、"openai" 或 "openai_compatible"。评判模型从 api_key_env 指定的环境变量中读取密钥(默认为 ANTHROPIC_API_KEY 或 OPENAI_API_KEY),并由该提供方计费。Groundedness 和 CitationSupport 通过 **options 接受 RubricJudge 的其余设置。概率评判模型、模型分类器和级联没有 SDK 类;它们只存在于 YAML 中。

ConversationCompleted 和 ConversationJudge 读取你的系统记录的 conversation/v1 产物。Oloproof 不驱动对话:你的应用运行每一轮并记录对话记录。见智能体。

@evaluator

def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

把一个单参数(即用例)的函数包装成一个 CustomEvaluator。用例具有 output、expected 和 scenario,而 artifacts(name) 返回某个已记录类型的载荷。函数可以是 def 或 async def。二值评估器返回 True 或 False;分数评估器声明 value_type="score" 和 score_range=(low, high),并返回一个数字。函数抛出的异常会把该用例在这个判据上记为缺失,而绝不会记为失败。

reads 必须列出函数读取的每个字段(input、output、expected、metadata、metadata.<key> 或 artifacts.<name>),因为缓存的评判结果正是以这些字段为键的。只有在 cacheable=True 时,评判结果才会在多次运行之间被复用。定义模块的源代码会计入版本,所以编辑它会使这些评判结果失效。YAML 无法指定自定义评估器。

诊断

diagnose 与 adiagnose

def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResult

在一种干预下重新执行已存储运行中失败的用例:"gold-context"、"top-k"(配合 top_k)或 "reranker"(配合 reranker)。system 和 evaluators 必须是该运行所用的版本;不同的版本会在任何执行之前被拒绝。使用 control=True 时,会在干预旁边运行一个全新的对照样本,从而把变更与运行之间的波动区分开。diagnose 是同步形式,在运行中的循环内部的行为与 evaluate 相同。工作流程见 RAG。

InterventionResult 保存父运行 id、干预、它是否受支持、干预运行和对照运行、每个用例的结果,以及在生成了报告时由 CaseDiagnosis 条目组成的 DiagnosisReport。

智能体重放

def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]

重放是你的系统做的事,而不是 Oloproof 模拟的事。系统只有通过实现 async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory,返回智能体从检查点开始所做的事,才能支持重放;Oloproof 把记录下来的前缀拼接上去并进行比较。你的应用负责其状态、会话和工具副作用,包括在重放之前重置它们。supports_replay 报告一个系统是否声明了该方法。

replay_case 是一个协程:await 它,或通过 asyncio.run 调用它。它在重放之前运行一个对照(使用 ReplayChange(kind="resume") 的同一个检查点),当对照无法复现记录时,不会尝试重放。change 是 ReplayChange(kind="drop_step", step_index=...) 或 ReplayChange(kind="resume")。checkpoint_for 选取严格位于某一步之前的最新记录检查点,或者 None,在后一种情况下,结果会以 no_checkpoint_recorded 被丢弃,而不触及系统。

ReplayOutcome 带有 scenario_id、change、重放和对照的终止状态,以及当该用例没有产生证据时带原因的 discarded。被丢弃的用例不是失败的用例。label_case 把一个结果转换为带有 FailureLabel 及其 LabelReason 的 CaseDiagnosis;unnecessary_steps 列出去掉后结果不变的步骤。见智能体。

人工标签与评估器信任

def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabel

存储一个人对已存储运行中某个用例给出的通过/失败结论。标签供 oloproof evaluators validate 使用,它度量评判模型与标签的一致性(AgreementResult)并记录其状态(RegistryEntry)。其余的 measurement_sample_* 参数把标签绑定到一个度量样本;CLI 的 oloproof labels export 和 oloproof labels import 会为你填写它们。见评判器。

在代码中编写策略

ReleasePolicy、IntervalThresholdRule、ObservedCountRule 和 DecisionRule(前两者的联合)无需文件即可构建策略。它们的字段就是配置参考中 release.yaml 的字段;区间规则使用 direction(min 或 max)和 threshold,而不是 min: 或 max:。比较规则没有导出的类;把它们写在 release.yaml 中并传入其路径。

from oloproof import IntervalThresholdRule, ReleasePolicy

policy = ReleasePolicy(
    rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)

其他导出

名称它是什么
ConcurrencyConfigsystem 和 judge 的限制,与 oloproof.yaml 中相同。
TransientError在系统中带 retryable=True 抛出它,让调用以退避方式重试。
ArtifactRef记录产物时返回的引用:它的类型和摘要。
__version__已安装的软件包版本。