指南
教學:分類器或結構化輸出
評估一個為客服問題打標籤的 Python 函式,讀懂發布為何被阻止,修復遺漏,並把修復與原版進行比較,全部在你自己的機器上完成,無需帳號、無需聯網,也無需任何模型。
你將建置什麼
一個客服機器人,返回一個包含 answer 和 label(refund、account 或 other)的 JSON 物件。你將用三條要求來衡量它:標籤足夠經常是正確的,輸出總是具有正確的結構,並且沒有任何回答洩露看起來像美國社會安全號碼的內容。其中兩條是不需要參考答案的格式檢查;一條依據參考標籤量測任務是否成功。兩者的區別很重要,本頁會把它們分開。
下文使用的術語(案例、執行、指標、區間、規則、閘門)在核心概念中有定義。
前提條件
- Python 3.11 或更高版本。
- 安裝在虛擬環境中的 Oloproof:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof- 範例專案和候選變更,它們隨軟體包一起提供。把兩者複製到新目錄中,並在第一個目錄裡工作;每個檔案也都列在下面,所以你也可以自己輸入:
oloproof init --example support_bot support-classifier
oloproof init --example classification support-change
cd support-classifier本頁任何地方都不使用 API 密鑰、提供者帳號或網路訪問。
檔案
support-classifier/
app.py the application under test (a Python callable)
oloproof.yaml the suite: dataset, system, evaluators
release.yaml the release policy: rules the run is decided against
data/support.jsonl 18 cases, one JSON object per line
rubrics/helpful.md a judge rubric, unused here在 support-classifier/ 目錄中執行每條命令。Oloproof 把它的儲存保存在那裡的 .oloproof/ 中;刪除該目錄即可從零重新開始。
應用及其適配器
你的應用通過適配器來訪問。對於 Python 應用,適配器就是函式本身:Oloproof 導入它,對每個案例用該案例的 input 呼叫它一次,並把它返回的字典記錄為該案例的輸出。
# app.py
from typing import Any
from oloproof import system
@system(name="support-bot", version="slice-a-example")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if "refund" in question:
return {"answer": "Refunds are available within 30 days when the order is eligible.",
"label": "refund"}
if "password" in question or "login" in question:
return {"answer": "Use password reset, then contact support if the login still fails.",
"label": "account"}
return {"answer": "A support specialist will follow up with the next step.", "label": "other"}要評估你自己的分類器,把它的代碼留在原處,並像這樣寫一個薄函式來呼叫它並返回一個字典。這個函式可以是 async。Oloproof 呼叫它;它不託管、不隔離也不重置你的應用,所以你的應用在兩次呼叫之間保存的任何狀態都由你自己管理。
oloproof.yaml 指明這個函式和評估器:
version: 1
project: support-bot-example
dataset: data/support.jsonl
system:
name: support-bot
version: slice-a-example
callable: app:answer
timeout_s: 30
evaluators:
- type: exact_match
criterion: exact_label
field: label
- type: json_schema
criterion: format_valid
field: null
schema:
type: object
required: [answer, label]
properties:
answer: {type: string}
label: {type: string}
additionalProperties: false
- type: regex
criterion: pii_free
field: answer
pattern: '\b\d{3}-\d{2}-\d{4}\b'
pass_if: no_match輸出按函式的原始碼、宣告的 version 和 config 快取。如果函式會讀取其他檔案(一個提示詞、一張規則表),把它們列在 system.code_paths 下,這樣編輯它們就會重新執行系統。
資料集
每行一個案例。input 正是你的函式作為 case 接收到的內容;expected 是 exact_match 評估器用來比較的參考:
{"id":"refund_00","input":{"question":"Can I get a refund for yesterday's order?"},"expected":{"label":"refund"}}
{"id":"account_04","input":{"question":"I can't sign in on my new phone."},"expected":{"label":"account"}}
{"id":"other_04","input":{"question":"I don't want a refund, I just need a copy of my receipt."},"expected":{"label":"other"}}對於每個案例,函式返回一個物件,例如 {"answer": "Use password reset, ...", "label": "account"}。
選擇評估器
| 判據 | 評估器 | 需要 expected | 它量測什麼 |
|---|---|---|---|
| exact_label | 對 label 使用 exact_match | 是 | 任務成功:標籤是正確的那個 |
| format_valid | 對整個輸出使用 json_schema | 否 | 格式:物件恰好有兩個字串欄位 |
| pii_free | 對 answer 使用 regex,pass_if: no_match | 否 | 文字的一項安全屬性 |
格式檢查會讓一個格式良好但錯誤的回答通過,所以它永遠不能替代任務成功。任務檢查需要每個案例都有參考;某個案例沒有參考時,exact_match 無法為它打分。確定性評估器不需要對照人工進行驗證:執行兩次會給出相同的結論。
政策
release.yaml 是這次執行據以決策的依據:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
rules:
- id: exact-label-floor
metric: exact_label
min: 0.70
- id: valid-format
metric: format_valid
kind: observed_count
max_failures: 0
- id: pii-free
metric: pii_free
kind: observed_count
max_failures: 0exact-label-floor 表示標籤至少在 70% 的情況下必須正確,並且只有當整個 95% 區間都不低於 0.70 時才通過。兩條 observed_count 規則在你執行的案例上完全不允許失敗;它們描述的是這些案例,而不是使用者將會提出的每一個問題。
執行
oloproof run真實輸出,有刪節:
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ exact-label-floor │ exact_label │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ valid-format │ format_valid │ PASS │ observed_failures_within_limit │
│ pii-free │ pii_free │ PASS │ observed_failures_within_limit │
│ exact_label │ 72.2% │ [46.5%, 90.4%] │ 13 / 18 observed · 0 missing · 0 excluded │
│ format_valid │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
│ pii_free │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/18 miss; judgment 0 hit/54 miss如何閱讀:
- 18 個標籤中有 13 個正確,即 72.2%。這高於 0.70,但區間向下一直延伸到 46.5%:18 個案例無法證明真實比率至少為 0.70。所以該規則是 INSUFFICIENT_EVIDENCE,既不是 PASS,也不是 FAIL。
- 每個輸出都具有正確的結構,並且沒有一個包含類似 SSN 的號碼,所以兩條格式規則都通過。
- block_on 列出了 INSUFFICIENT_EVIDENCE,所以閘門阻止發布,命令以 3 結束。結束代碼 0 表示沒有出現任何政策要阻止的情況;CI 閘門列出了所有結束代碼。
再執行一次,快取行顯示為 execution 18 hit/0 miss:沒有任何變化,所以函式不會被呼叫。
檢視失敗項
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_04RUN_ID 是執行輸出第一行中的 id。
5 of 18 cases failed, errored or did not finish
refund_04
output: {"answer": "A support specialist will follow up with the next step.", "label": "other"}
exact_label: failed
...
other_04
output: {"answer": "Refunds are available within 30 days when the order is eligible.", "label": "refund"}
exact_label: failedcase refund_04
input: {
"question": "I was charged twice this month and want my money back."
}
expected: {
"label": "refund"
}
execution: OK, 1 ms
output: {
"answer": "A support specialist will follow up with the next step.",
"label": "other"
}
judgments:
exact_label: failed
format_valid: passed
pii_free: passed一旦讀了輸入,規律就很明顯:“money back”、“reverse the payment”、“sign in”和“two-factor”都不在關鍵詞清單中,而 other_04 說的是“I don't want a refund”,單詞“refund”照樣匹配上了。注意 refund_04 雖然是錯的,卻通過了兩項格式檢查:這正是檢查格式與量測成功之間的差距。
這裡有兩個有意義的下一步。修復遺漏(見下文),或者增加案例:在相同準確率下案例越多,區間越窄,oloproof plan RUN_ID --run 會估計需要多少案例。
做一次真正的修改
把 ../support-change/app.py 複製到 app.py 上覆蓋它。它加入了被遺漏的說法:
REFUND_WORDS = ("refund", "money back", "reverse the payment")
ACCOUNT_WORDS = ("password", "login", "sign in", "two-factor")
@system(name="support-bot", version="keywords-v2")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if any(word in question for word in REFUND_WORDS):
...並在 oloproof.yaml 的 system 下設置 version: keywords-v2,讓這次執行被記錄為新版本。然後:
oloproof runGate: ALLOW (exit 0)
│ exact-label-floor │ exact_label │ PASS │ lower_bound_meets_minimum │
│ exact_label │ 94.4% │ [72.7%, 99.9%] │ 17 / 18 observed · 0 missing · 0 excluded │18 個中有 17 個正確,區間下界 72.7% 越過了 0.70,所以該規則通過,命令以 0 結束。other_04 仍然失敗:這次修復沒有處理否定。
把候選與基準進行比較
執行規則問的是候選是否達到你的下限。比較問的是它與基準逐個案例有何不同。把 ../support-change/compare.yaml 複製到專案中;它包含一條比較規則:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: label-no-regression
kind: non_inferiority
metric: exact_label
margin: 0.10oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlComparison sha256:de76... of run_01M4...TJAD against run_01M4...ECVEF · 18 paired cases
exact_label: +22.2 points [-12.9, +57.0] · 18 paired · 0 missing · 0 excluded
format_valid: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
pii_free: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
Decisions
label-no-regression exact_label non-inferiority, margin 10.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 3 more paired cases would decide it, if the difference holds (21 in total at 22% discordance)
Gate: BLOCK (exit 3)候選修復了四個案例,沒有弄壞任何一個,估計提升 22 個百分點。但只有四個案例發生了變化,18 個配對案例留下的區間從差 12.9 個百分點到好 57 個百分點,跨越了 10 個百分點的邊際。這次比較尚不能排除候選比你所能接受的程度更差,所以結果是 INSUFFICIENT_EVIDENCE,以 3 結束。它下面那一行是樣本數估算。不帶 --policy 時,compare 會印出差異,說明專案的 release.yaml 沒有宣告任何比較規則,並以 0 結束,因為什麼都沒有被決策。
將候選與基準進行比較解釋了邊際以及其他規則類型。
故障排除
| 症狀 | 原因與修復 |
|---|---|
| app 的 ModuleNotFoundError | 在包含 app.py 的目錄中執行,或給 callable 一個可以從那裡導入的模組路徑。 |
| 規則指定了一個沒有任何評估器產出的指標 | 規則的 metric 必須等於某個評估器的 criterion;報錯會列出現有的指標。 |
| 你編輯了分類器,但執行復用了所有輸出 | 快取跟隨可呼叫物件的原始碼;它讀取的輔助檔案必須列在 system.code_paths 下。 |
| exact_label 報告有案例缺失 | 這些執行拋出了異常或超時;oloproof inspect RUN_ID --failures 會顯示每個錯誤。 |
| 估計值很高,執行卻以 3 結束 | 起決定作用的是區間,而不是估計值。增加案例,或接受一個更低的下限,並在執行之前決定。 |