指南
SDK 参考
oloproof 和 oloproof.evaluators 包导出的每个名称,附带其签名、它是同步还是异步的,以及它返回什么。如需引导式介绍,请先阅读 Python API。
只有这两个包是公开接口。从 oloproof_core 导入的任何东西都是引擎内部实现,可能在不另行通知的情况下改变。下面的每个函数都在本地针对项目的存储运行;除非你传入的某个评估器调用了模型提供方,否则它们都不会把数据发送到任何地方。
运行评估
evaluate 与 aevaluate
def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult| 参数 | 类型 | 它是什么 |
|---|---|---|
| system | 一个 @system 函数、一个 @rag_system 类或实例,或一个可调用对象 | 被测系统。 |
| dataset | 路径 | 一个 JSONL 测试套件。见测试套件。 |
| evaluators | 列表 | 来自 oloproof.evaluators 的实例,或 @evaluator 函数。 |
| policy | release.yaml 的路径、一个 ReleasePolicy,或 None | 发布策略。None 不运行门禁:result.gate 为 None,什么都不决策。 |
| concurrency | ConcurrencyConfig 或像 {"system": 8, "judge": 4} 这样的映射 | 同时进行的调用数。 |
| slices | 字符串列表 | 探索性切片,与 oloproof.yaml 中相同。 |
| min_slice_support | 整数 | 符合条件的用例少于这个数时,切片没有区间。默认 30。 |
| replicates | 整数 | 每个用例度量这么多次。默认 1。 |
evaluate 是同步的。在没有运行中的事件循环时调用,它使用 asyncio.run;在运行中的循环内部调用时(笔记本、异步测试),它在一个单独的线程上运行评估并阻塞直到完成,所以在两种场合都可以安全使用。aevaluate 是协程;在异步代码中 await 它。
其余的关键字参数(metrics、store、predictive、event_sink、retry_policy、traffic_draw_id)接受来自 oloproof_core 的引擎类型,不属于稳定接口。
from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator
@system(name="support-bot", version="1")
def answer(case):
current_case().usage(input_tokens=12, output_tokens=3)
return {"label": "refund" if "refund" in case["question"].lower() else "other"}
@evaluator(criterion="short_label")
def short_label(case):
return len(case.output["label"]) <= 6
result = evaluate(
system=answer,
dataset="cases.jsonl",
evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)在一个两用例的测试套件上不带策略运行,打印出:
correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0EvaluationResult
| 成员 | 类型 | 它是什么 |
|---|---|---|
| run | 运行记录 | 已存储的运行,带有它的 id、状态和完整性。 |
| suite、system、evaluators | 版本记录 | 这次运行所度量的确切版本。 |
| metrics | 指标结果的元组 | 每个判据和声明的指标一个:metric、estimate、interval、n_total、n_eligible、n_observed、n_missing、exclusions、method。 |
| gate | 门禁结果或 None | 有策略时:release_action、exit_code、decisions 和 reasons。 |
| decisions | 元组 | 门禁的决策,没有策略时为空。 |
| cases() | 列表 | 每个用例及其执行和评判结果。 |
| failures() | 列表 | 没有完成、出错或至少在一个评估器上失败的用例。 |
| print(stderr=False) | 无 | oloproof run 打印的终端报告。 |
| to_bundle(path) | 路径 | 写出一个可移植的包,与 oloproof export 相同。 |
状态、原因代码和计数的含义见结果与执行。
evaluate_comparison 与 aevaluate_comparison
def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult以同一个带种子的顺序在同一个测试套件上运行两个系统,并对策略中的比较规则作出决策。policy 是必需的,并且必须至少包含一条 superiority、non_inferiority 或 equivalence 规则,否则调用会抛出配置错误。额外的关键字参数:concurrency、replicates,以及引擎类型的 metrics、store 和 retry_policy。同步和异步的行为与 evaluate 相同。
ComparisonEvaluationResult 包含 candidate 和 baseline(各为一个 EvaluationResult)以及 comparison,后者带有配对差异及其决策。见比较两个版本。
声明系统
@system
def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())可以直接使用(@system),也可以带参数使用(@system(name=..., version=...)),或者对一个对象调用(system(model.answer, version="v2"))。函数接收的是用例的 input 对象,而不是整个用例,并返回评估器读取的输出。它可以是 def 或 async def;同步函数在工作线程上运行。
| 参数 | 默认值 | 它是什么 |
|---|---|---|
| name | 函数的名称 | 版本身份的一部分。 |
| version | 无 | 对于绑定方法或可调用对象是必需的,因为它们的行为取决于 Oloproof 看不到的状态。 |
| config | 空 | 与版本一同记录的设置。 |
| timeout_s | 120 | 每次调用的时限。超时的调用会被记录为超时的执行。 |
| records | 空 | 系统记录的产物类型,例如 retrieval/v1。如果某个评估器需要一个系统没有声明的类型,它会在运行开始之前被拒绝。 |
函数自身模块的源代码会计入版本摘要,所以编辑它会使缓存的执行失效。还有哪些会、哪些不会,见配置参考。
current_case
def current_case() -> CaseRecorder只有在 Oloproof 调用你的系统期间才可用;在其他任何地方它都会抛出 RuntimeError。记录器的方法:
| 方法 | 记录 |
|---|---|
| usage(*, input_tokens=None, output_tokens=None, cost_usd=None) | 一次模型调用的 token 数和成本。省略的值保持未记录,而不是零。 |
| artifact(kind, data) | 任意 JSON 值或 Pydantic 模型,归入 trace 或 conversation/v1 这样的类型下。 |
| retrieval(retrieval) | 检索器返回的排序后的候选(retrieval/v1)。 |
| context(context) | 为生成组装的上下文(context/v1)。 |
| citations(ids) | 回答引用的 id,形式为 doc_id 或 doc_id#chunk_id(citations/v1)。 |
| agent_trajectory(trajectory) | 智能体的步骤、工具调用和结果,以及检查点(agent_trajectory/v1)。 |
每个方法都返回一个 ArtifactRef(usage 除外,它什么都不返回)。为了构建这些记录,导出了以下有类型的载荷:Retrieval、Passage、Context、ContextItem、DroppedItem、Citations、StageTimings、AgentTrajectory、AgentStep、AgentCheckpoint、AgentConstraintCheck,以及类型名 CONVERSATION(conversation/v1)。
@rag_system
def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")一个类装饰器。该类提供 retrieve(input, depth) 和 generate(input, context),设置了 token_budget 时还要提供 count_tokens(passage)。context 是经过 top_k 和预算筛选后留下的 Passage 对象列表,按排名顺序排列。Oloproof 自己记录 retrieval/v1、context/v1、citations/v1 和 stage_timings/v1,并分别缓存每个阶段。citations_path 指明保存回答所引用 id 的输出字段。见 RAG。
评估器
所有类都在 oloproof.evaluators 中。每个类的 criterion 指明它产生的指标。每个评估器读取哪些产物,以及它在 YAML 中的对应写法,见配置参考中的评估器表。
| 类 | 签名 |
|---|---|
| ExactMatch | (*, criterion, field=None, expected_field=None, strip=True, casefold=False) |
| Contains | (*, criterion, field=None, expected_field=None) |
| Regex | (*, criterion, pattern, field=None, pass_if="match") |
| JsonSchema | (*, criterion, schema, field=None) |
| RubricJudge | (*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0) |
| Groundedness | (*, provider, model, criterion="groundedness", **options) |
| CitationSupport | (*, provider, model, criterion="citation_support", **options) |
| CitationValidity | (*, criterion="citations_valid", require_citations=False) |
| HitRate、Recall | (k=None, *, criterion=None, relevance_unit="doc"),k 默认为 5 |
| MRR、NDCG | (k=None, *, criterion=None, relevance_unit="doc"),k 默认为 10 |
| AgentMaxSteps | (max_steps, *, criterion=None) |
| AgentToolCalled | (tool_name, *, min_calls=1, criterion=None) |
| AgentNoToolLoop | (*, max_repeats=2, criterion="agent_no_tool_loop") |
| AgentToolSequence | (*, ordered=True, criterion="agent_tool_sequence") |
| AgentNoUndeclaredTool | (*, criterion="agent_no_undeclared_tool") |
| AgentConstraintsSatisfied | (constraints=(), *, criterion="agent_constraints_satisfied") |
| AgentRoute | (*, criterion="agent_route") |
| AgentToolPermissions | (permissions, *, criterion="agent_tool_permissions") |
| AgentMaxHandoffs | (max_handoffs, *, criterion=None) |
| ConversationCompleted | (*, criterion="conversation_completed") |
| ConversationJudge | 与 RubricJudge 相同 |
| PredictiveCorrect、PredictiveRecall、PredictivePrecision | (*, criterion, positive=True, field="label", expected_field="label") |
| AbsoluteError | (*, criterion, target_range, field="label", expected_field="label") |
| Brier、PredictiveRanking | (*, criterion, positive=True, field="score", expected_field="label") |
| LogLoss | (*, clip, criterion, positive=True, field="score", expected_field="label") |
| CustomEvaluator | (func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None) |
provider 是 "anthropic"、"openai" 或 "openai_compatible"。评判模型从 api_key_env 指定的环境变量中读取密钥(默认为 ANTHROPIC_API_KEY 或 OPENAI_API_KEY),并由该提供方计费。Groundedness 和 CitationSupport 通过 **options 接受 RubricJudge 的其余设置。概率评判模型、模型分类器和级联没有 SDK 类;它们只存在于 YAML 中。
ConversationCompleted 和 ConversationJudge 读取你的系统记录的 conversation/v1 产物。Oloproof 不驱动对话:你的应用运行每一轮并记录对话记录。见智能体。
@evaluator
def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)把一个单参数(即用例)的函数包装成一个 CustomEvaluator。用例具有 output、expected 和 scenario,而 artifacts(name) 返回某个已记录类型的载荷。函数可以是 def 或 async def。二值评估器返回 True 或 False;分数评估器声明 value_type="score" 和 score_range=(low, high),并返回一个数字。函数抛出的异常会把该用例在这个判据上记为缺失,而绝不会记为失败。
reads 必须列出函数读取的每个字段(input、output、expected、metadata、metadata.<key> 或 artifacts.<name>),因为缓存的评判结果正是以这些字段为键的。只有在 cacheable=True 时,评判结果才会在多次运行之间被复用。定义模块的源代码会计入版本,所以编辑它会使这些评判结果失效。YAML 无法指定自定义评估器。
诊断
diagnose 与 adiagnose
def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResult在一种干预下重新执行已存储运行中失败的用例:"gold-context"、"top-k"(配合 top_k)或 "reranker"(配合 reranker)。system 和 evaluators 必须是该运行所用的版本;不同的版本会在任何执行之前被拒绝。使用 control=True 时,会在干预旁边运行一个全新的对照样本,从而把变更与运行之间的波动区分开。diagnose 是同步形式,在运行中的循环内部的行为与 evaluate 相同。工作流程见 RAG。
InterventionResult 保存父运行 id、干预、它是否受支持、干预运行和对照运行、每个用例的结果,以及在生成了报告时由 CaseDiagnosis 条目组成的 DiagnosisReport。
智能体重放
def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]重放是你的系统做的事,而不是 Oloproof 模拟的事。系统只有通过实现 async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory,返回智能体从检查点开始所做的事,才能支持重放;Oloproof 把记录下来的前缀拼接上去并进行比较。你的应用负责其状态、会话和工具副作用,包括在重放之前重置它们。supports_replay 报告一个系统是否声明了该方法。
replay_case 是一个协程:await 它,或通过 asyncio.run 调用它。它在重放之前运行一个对照(使用 ReplayChange(kind="resume") 的同一个检查点),当对照无法复现记录时,不会尝试重放。change 是 ReplayChange(kind="drop_step", step_index=...) 或 ReplayChange(kind="resume")。checkpoint_for 选取严格位于某一步之前的最新记录检查点,或者 None,在后一种情况下,结果会以 no_checkpoint_recorded 被丢弃,而不触及系统。
ReplayOutcome 带有 scenario_id、change、重放和对照的终止状态,以及当该用例没有产生证据时带原因的 discarded。被丢弃的用例不是失败的用例。label_case 把一个结果转换为带有 FailureLabel 及其 LabelReason 的 CaseDiagnosis;unnecessary_steps 列出去掉后结果不变的步骤。见智能体。
人工标签与评估器信任
def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabel存储一个人对已存储运行中某个用例给出的通过/失败结论。标签供 oloproof evaluators validate 使用,它度量评判模型与标签的一致性(AgreementResult)并记录其状态(RegistryEntry)。其余的 measurement_sample_* 参数把标签绑定到一个度量样本;CLI 的 oloproof labels export 和 oloproof labels import 会为你填写它们。见评判器。
在代码中编写策略
ReleasePolicy、IntervalThresholdRule、ObservedCountRule 和 DecisionRule(前两者的联合)无需文件即可构建策略。它们的字段就是配置参考中 release.yaml 的字段;区间规则使用 direction(min 或 max)和 threshold,而不是 min: 或 max:。比较规则没有导出的类;把它们写在 release.yaml 中并传入其路径。
from oloproof import IntervalThresholdRule, ReleasePolicy
policy = ReleasePolicy(
rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)其他导出
| 名称 | 它是什么 |
|---|---|
| ConcurrencyConfig | system 和 judge 的限制,与 oloproof.yaml 中相同。 |
| TransientError | 在系统中带 retryable=True 抛出它,让调用以退避方式重试。 |
| ArtifactRef | 记录产物时返回的引用:它的类型和摘要。 |
| __version__ | 已安装的软件包版本。 |