指南
教程:分类器或结构化输出
评估一个为客服问题打标签的 Python 函数,读懂发布为何被阻止,修复遗漏,并把修复与原版进行比较,全部在你自己的机器上完成,无需账户、无需联网,也无需任何模型。
你将构建什么
一个客服机器人,返回一个包含 answer 和 label(refund、account 或 other)的 JSON 对象。你将用三条要求来衡量它:标签足够经常是正确的,输出总是具有正确的结构,并且没有任何回答泄露看起来像美国社会安全号码的内容。其中两条是不需要参考答案的格式检查;一条依据参考标签度量任务是否成功。两者的区别很重要,本页会把它们分开。
下文使用的术语(用例、运行、指标、区间、规则、门禁)在核心概念中有定义。
前提条件
- Python 3.11 或更高版本。
- 安装在虚拟环境中的 Oloproof:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof- 示例项目和候选变更,它们随软件包一起提供。把两者复制到新目录中,并在第一个目录里工作;每个文件也都列在下面,所以你也可以自己输入:
oloproof init --example support_bot support-classifier
oloproof init --example classification support-change
cd support-classifier本页任何地方都不使用 API 密钥、提供方账户或网络访问。
文件
support-classifier/
app.py the application under test (a Python callable)
oloproof.yaml the suite: dataset, system, evaluators
release.yaml the release policy: rules the run is decided against
data/support.jsonl 18 cases, one JSON object per line
rubrics/helpful.md a judge rubric, unused here在 support-classifier/ 目录中运行每条命令。Oloproof 把它的存储保存在那里的 .oloproof/ 中;删除该目录即可从零重新开始。
应用及其适配器
你的应用通过适配器来访问。对于 Python 应用,适配器就是函数本身:Oloproof 导入它,对每个用例用该用例的 input 调用它一次,并把它返回的字典记录为该用例的输出。
# app.py
from typing import Any
from oloproof import system
@system(name="support-bot", version="slice-a-example")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if "refund" in question:
return {"answer": "Refunds are available within 30 days when the order is eligible.",
"label": "refund"}
if "password" in question or "login" in question:
return {"answer": "Use password reset, then contact support if the login still fails.",
"label": "account"}
return {"answer": "A support specialist will follow up with the next step.", "label": "other"}要评估你自己的分类器,把它的代码留在原处,并像这样写一个薄函数来调用它并返回一个字典。这个函数可以是 async。Oloproof 调用它;它不托管、不隔离也不重置你的应用,所以你的应用在两次调用之间保存的任何状态都由你自己管理。
oloproof.yaml 指明这个函数和评估器:
version: 1
project: support-bot-example
dataset: data/support.jsonl
system:
name: support-bot
version: slice-a-example
callable: app:answer
timeout_s: 30
evaluators:
- type: exact_match
criterion: exact_label
field: label
- type: json_schema
criterion: format_valid
field: null
schema:
type: object
required: [answer, label]
properties:
answer: {type: string}
label: {type: string}
additionalProperties: false
- type: regex
criterion: pii_free
field: answer
pattern: '\b\d{3}-\d{2}-\d{4}\b'
pass_if: no_match输出按函数的源代码、声明的 version 和 config 缓存。如果函数会读取其他文件(一个提示词、一张规则表),把它们列在 system.code_paths 下,这样编辑它们就会重新运行系统。
数据集
每行一个用例。input 正是你的函数作为 case 接收到的内容;expected 是 exact_match 评估器用来比较的参考:
{"id":"refund_00","input":{"question":"Can I get a refund for yesterday's order?"},"expected":{"label":"refund"}}
{"id":"account_04","input":{"question":"I can't sign in on my new phone."},"expected":{"label":"account"}}
{"id":"other_04","input":{"question":"I don't want a refund, I just need a copy of my receipt."},"expected":{"label":"other"}}对于每个用例,函数返回一个对象,例如 {"answer": "Use password reset, ...", "label": "account"}。
选择评估器
| 判据 | 评估器 | 需要 expected | 它度量什么 |
|---|---|---|---|
| exact_label | 对 label 使用 exact_match | 是 | 任务成功:标签是正确的那个 |
| format_valid | 对整个输出使用 json_schema | 否 | 格式:对象恰好有两个字符串字段 |
| pii_free | 对 answer 使用 regex,pass_if: no_match | 否 | 文本的一项安全属性 |
格式检查会让一个格式良好但错误的回答通过,所以它永远不能替代任务成功。任务检查需要每个用例都有参考;某个用例没有参考时,exact_match 无法为它打分。确定性评估器不需要对照人工进行验证:运行两次会给出相同的结论。
策略
release.yaml 是这次运行据以决策的依据:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
rules:
- id: exact-label-floor
metric: exact_label
min: 0.70
- id: valid-format
metric: format_valid
kind: observed_count
max_failures: 0
- id: pii-free
metric: pii_free
kind: observed_count
max_failures: 0exact-label-floor 表示标签至少在 70% 的情况下必须正确,并且只有当整个 95% 区间都不低于 0.70 时才通过。两条 observed_count 规则在你运行的用例上完全不允许失败;它们描述的是这些用例,而不是用户将会提出的每一个问题。
运行
oloproof run真实输出,有删节:
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ exact-label-floor │ exact_label │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ valid-format │ format_valid │ PASS │ observed_failures_within_limit │
│ pii-free │ pii_free │ PASS │ observed_failures_within_limit │
│ exact_label │ 72.2% │ [46.5%, 90.4%] │ 13 / 18 observed · 0 missing · 0 excluded │
│ format_valid │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
│ pii_free │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/18 miss; judgment 0 hit/54 miss如何阅读:
- 18 个标签中有 13 个正确,即 72.2%。这高于 0.70,但区间向下一直延伸到 46.5%:18 个用例无法证明真实比率至少为 0.70。所以该规则是 INSUFFICIENT_EVIDENCE,既不是 PASS,也不是 FAIL。
- 每个输出都具有正确的结构,并且没有一个包含类似 SSN 的号码,所以两条格式规则都通过。
- block_on 列出了 INSUFFICIENT_EVIDENCE,所以门禁阻止发布,命令以 3 退出。退出码 0 表示没有出现任何策略要阻止的情况;CI 门禁列出了所有退出码。
再运行一次,缓存行显示为 execution 18 hit/0 miss:没有任何变化,所以函数不会被调用。
查看失败项
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_04RUN_ID 是运行输出第一行中的 id。
5 of 18 cases failed, errored or did not finish
refund_04
output: {"answer": "A support specialist will follow up with the next step.", "label": "other"}
exact_label: failed
...
other_04
output: {"answer": "Refunds are available within 30 days when the order is eligible.", "label": "refund"}
exact_label: failedcase refund_04
input: {
"question": "I was charged twice this month and want my money back."
}
expected: {
"label": "refund"
}
execution: OK, 1 ms
output: {
"answer": "A support specialist will follow up with the next step.",
"label": "other"
}
judgments:
exact_label: failed
format_valid: passed
pii_free: passed一旦读了输入,规律就很明显:“money back”、“reverse the payment”、“sign in”和“two-factor”都不在关键词列表中,而 other_04 说的是“I don't want a refund”,单词“refund”照样匹配上了。注意 refund_04 虽然是错的,却通过了两项格式检查:这正是检查格式与度量成功之间的差距。
这里有两个有意义的下一步。修复遗漏(见下文),或者增加用例:在相同准确率下用例越多,区间越窄,oloproof plan RUN_ID --run 会估计需要多少用例。
做一次真正的修改
把 ../support-change/app.py 复制到 app.py 上覆盖它。它加入了被遗漏的说法:
REFUND_WORDS = ("refund", "money back", "reverse the payment")
ACCOUNT_WORDS = ("password", "login", "sign in", "two-factor")
@system(name="support-bot", version="keywords-v2")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if any(word in question for word in REFUND_WORDS):
...并在 oloproof.yaml 的 system 下设置 version: keywords-v2,让这次运行被记录为新版本。然后:
oloproof runGate: ALLOW (exit 0)
│ exact-label-floor │ exact_label │ PASS │ lower_bound_meets_minimum │
│ exact_label │ 94.4% │ [72.7%, 99.9%] │ 17 / 18 observed · 0 missing · 0 excluded │18 个中有 17 个正确,区间下界 72.7% 越过了 0.70,所以该规则通过,命令以 0 退出。other_04 仍然失败:这次修复没有处理否定。
把候选与基线进行比较
运行规则问的是候选是否达到你的下限。比较问的是它与基线逐个用例有何不同。把 ../support-change/compare.yaml 复制到项目中;它包含一条比较规则:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: label-no-regression
kind: non_inferiority
metric: exact_label
margin: 0.10oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlComparison sha256:de76... of run_01M4...TJAD against run_01M4...ECVEF · 18 paired cases
exact_label: +22.2 points [-12.9, +57.0] · 18 paired · 0 missing · 0 excluded
format_valid: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
pii_free: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
Decisions
label-no-regression exact_label non-inferiority, margin 10.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 3 more paired cases would decide it, if the difference holds (21 in total at 22% discordance)
Gate: BLOCK (exit 3)候选修复了四个用例,没有弄坏任何一个,估计提升 22 个百分点。但只有四个用例发生了变化,18 个配对用例留下的区间从差 12.9 个百分点到好 57 个百分点,跨越了 10 个百分点的边际。这次比较尚不能排除候选比你所能接受的程度更差,所以结果是 INSUFFICIENT_EVIDENCE,以 3 退出。它下面那一行是样本量估算。不带 --policy 时,compare 会打印差异,说明项目的 release.yaml 没有声明任何比较规则,并以 0 退出,因为什么都没有被决策。
将候选与基线进行比较解释了边际以及其他规则类型。
故障排除
| 症状 | 原因与修复 |
|---|---|
| app 的 ModuleNotFoundError | 在包含 app.py 的目录中运行,或给 callable 一个可以从那里导入的模块路径。 |
| 规则指定了一个没有任何评估器产出的指标 | 规则的 metric 必须等于某个评估器的 criterion;报错会列出现有的指标。 |
| 你编辑了分类器,但运行复用了所有输出 | 缓存跟随可调用对象的源代码;它读取的辅助文件必须列在 system.code_paths 下。 |
| exact_label 报告有用例缺失 | 这些执行抛出了异常或超时;oloproof inspect RUN_ID --failures 会显示每个错误。 |
| 估计值很高,运行却以 3 退出 | 起决定作用的是区间,而不是估计值。增加用例,或接受一个更低的下限,并在运行之前决定。 |