跳至主要內容

指南

教學:分類器或結構化輸出

評估一個為客服問題打標籤的 Python 函式,讀懂發布為何被阻止,修復遺漏,並把修復與原版進行比較,全部在你自己的機器上完成,無需帳號、無需聯網,也無需任何模型。

你將建置什麼

一個客服機器人,返回一個包含 answer 和 label(refund、account 或 other)的 JSON 物件。你將用三條要求來衡量它:標籤足夠經常是正確的,輸出總是具有正確的結構,並且沒有任何回答洩露看起來像美國社會安全號碼的內容。其中兩條是不需要參考答案的格式檢查;一條依據參考標籤量測任務是否成功。兩者的區別很重要,本頁會把它們分開。

下文使用的術語(案例、執行、指標、區間、規則、閘門)在核心概念中有定義。

前提條件

  • Python 3.11 或更高版本。
  • 安裝在虛擬環境中的 Oloproof:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • 範例專案和候選變更,它們隨軟體包一起提供。把兩者複製到新目錄中,並在第一個目錄裡工作;每個檔案也都列在下面,所以你也可以自己輸入:
oloproof init --example support_bot support-classifier
oloproof init --example classification support-change
cd support-classifier

本頁任何地方都不使用 API 密鑰、提供者帳號或網路訪問。

檔案

support-classifier/
  app.py              the application under test (a Python callable)
  oloproof.yaml       the suite: dataset, system, evaluators
  release.yaml        the release policy: rules the run is decided against
  data/support.jsonl  18 cases, one JSON object per line
  rubrics/helpful.md  a judge rubric, unused here

在 support-classifier/ 目錄中執行每條命令。Oloproof 把它的儲存保存在那裡的 .oloproof/ 中;刪除該目錄即可從零重新開始。

應用及其適配器

你的應用通過適配器來訪問。對於 Python 應用,適配器就是函式本身:Oloproof 導入它,對每個案例用該案例的 input 呼叫它一次,並把它返回的字典記錄為該案例的輸出。

# app.py
from typing import Any

from oloproof import system


@system(name="support-bot", version="slice-a-example")
def answer(case: dict[str, Any]) -> dict[str, str]:
    question = str(case["question"]).lower()
    if "refund" in question:
        return {"answer": "Refunds are available within 30 days when the order is eligible.",
                "label": "refund"}
    if "password" in question or "login" in question:
        return {"answer": "Use password reset, then contact support if the login still fails.",
                "label": "account"}
    return {"answer": "A support specialist will follow up with the next step.", "label": "other"}

要評估你自己的分類器,把它的代碼留在原處,並像這樣寫一個薄函式來呼叫它並返回一個字典。這個函式可以是 async。Oloproof 呼叫它;它不託管、不隔離也不重置你的應用,所以你的應用在兩次呼叫之間保存的任何狀態都由你自己管理。

oloproof.yaml 指明這個函式和評估器:

version: 1
project: support-bot-example
dataset: data/support.jsonl
system:
  name: support-bot
  version: slice-a-example
  callable: app:answer
  timeout_s: 30
evaluators:
  - type: exact_match
    criterion: exact_label
    field: label
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [answer, label]
      properties:
        answer: {type: string}
        label: {type: string}
      additionalProperties: false
  - type: regex
    criterion: pii_free
    field: answer
    pattern: '\b\d{3}-\d{2}-\d{4}\b'
    pass_if: no_match

輸出按函式的原始碼、宣告的 version 和 config 快取。如果函式會讀取其他檔案(一個提示詞、一張規則表),把它們列在 system.code_paths 下,這樣編輯它們就會重新執行系統。

資料集

每行一個案例。input 正是你的函式作為 case 接收到的內容;expected 是 exact_match 評估器用來比較的參考:

{"id":"refund_00","input":{"question":"Can I get a refund for yesterday's order?"},"expected":{"label":"refund"}}
{"id":"account_04","input":{"question":"I can't sign in on my new phone."},"expected":{"label":"account"}}
{"id":"other_04","input":{"question":"I don't want a refund, I just need a copy of my receipt."},"expected":{"label":"other"}}

對於每個案例,函式返回一個物件,例如 {"answer": "Use password reset, ...", "label": "account"}。

選擇評估器

判據評估器需要 expected它量測什麼
exact_label對 label 使用 exact_match是任務成功:標籤是正確的那個
format_valid對整個輸出使用 json_schema否格式:物件恰好有兩個字串欄位
pii_free對 answer 使用 regex,pass_if: no_match否文字的一項安全屬性

格式檢查會讓一個格式良好但錯誤的回答通過,所以它永遠不能替代任務成功。任務檢查需要每個案例都有參考;某個案例沒有參考時,exact_match 無法為它打分。確定性評估器不需要對照人工進行驗證:執行兩次會給出相同的結論。

政策

release.yaml 是這次執行據以決策的依據:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
rules:
  - id: exact-label-floor
    metric: exact_label
    min: 0.70
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: pii-free
    metric: pii_free
    kind: observed_count
    max_failures: 0

exact-label-floor 表示標籤至少在 70% 的情況下必須正確,並且只有當整個 95% 區間都不低於 0.70 時才通過。兩條 observed_count 規則在你執行的案例上完全不允許失敗;它們描述的是這些案例,而不是使用者將會提出的每一個問題。

執行

oloproof run

真實輸出,有刪節:

Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ exact-label-floor │ exact_label  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ valid-format      │ format_valid │ PASS                  │ observed_failures_within_limit │
│ pii-free          │ pii_free     │ PASS                  │ observed_failures_within_limit │
│ exact_label  │ 72.2%    │ [46.5%, 90.4%]  │ 13 / 18 observed · 0 missing · 0 excluded │
│ format_valid │ 100.0%   │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
│ pii_free     │ 100.0%   │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/18 miss; judgment 0 hit/54 miss

如何閱讀:

  • 18 個標籤中有 13 個正確,即 72.2%。這高於 0.70,但區間向下一直延伸到 46.5%:18 個案例無法證明真實比率至少為 0.70。所以該規則是 INSUFFICIENT_EVIDENCE,既不是 PASS,也不是 FAIL。
  • 每個輸出都具有正確的結構,並且沒有一個包含類似 SSN 的號碼,所以兩條格式規則都通過。
  • block_on 列出了 INSUFFICIENT_EVIDENCE,所以閘門阻止發布,命令以 3 結束。結束代碼 0 表示沒有出現任何政策要阻止的情況;CI 閘門列出了所有結束代碼。

再執行一次,快取行顯示為 execution 18 hit/0 miss:沒有任何變化,所以函式不會被呼叫。

檢視失敗項

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_04

RUN_ID 是執行輸出第一行中的 id。

5 of 18 cases failed, errored or did not finish

refund_04
  output: {"answer": "A support specialist will follow up with the next step.", "label": "other"}
  exact_label: failed
...
other_04
  output: {"answer": "Refunds are available within 30 days when the order is eligible.", "label": "refund"}
  exact_label: failed
case refund_04
input: {
  "question": "I was charged twice this month and want my money back."
}
expected: {
  "label": "refund"
}
execution: OK, 1 ms
output: {
  "answer": "A support specialist will follow up with the next step.",
  "label": "other"
}
judgments:
  exact_label: failed
  format_valid: passed
  pii_free: passed

一旦讀了輸入,規律就很明顯:“money back”、“reverse the payment”、“sign in”和“two-factor”都不在關鍵詞清單中,而 other_04 說的是“I don't want a refund”,單詞“refund”照樣匹配上了。注意 refund_04 雖然是錯的,卻通過了兩項格式檢查:這正是檢查格式與量測成功之間的差距。

這裡有兩個有意義的下一步。修復遺漏(見下文),或者增加案例:在相同準確率下案例越多,區間越窄,oloproof plan RUN_ID --run 會估計需要多少案例。

做一次真正的修改

把 ../support-change/app.py 複製到 app.py 上覆蓋它。它加入了被遺漏的說法:

REFUND_WORDS = ("refund", "money back", "reverse the payment")
ACCOUNT_WORDS = ("password", "login", "sign in", "two-factor")


@system(name="support-bot", version="keywords-v2")
def answer(case: dict[str, Any]) -> dict[str, str]:
    question = str(case["question"]).lower()
    if any(word in question for word in REFUND_WORDS):
        ...

並在 oloproof.yaml 的 system 下設置 version: keywords-v2,讓這次執行被記錄為新版本。然後:

oloproof run
Gate: ALLOW (exit 0)
│ exact-label-floor │ exact_label  │ PASS  │ lower_bound_meets_minimum      │
│ exact_label  │ 94.4%    │ [72.7%, 99.9%]  │ 17 / 18 observed · 0 missing · 0 excluded │

18 個中有 17 個正確,區間下界 72.7% 越過了 0.70,所以該規則通過,命令以 0 結束。other_04 仍然失敗:這次修復沒有處理否定。

把候選與基準進行比較

執行規則問的是候選是否達到你的下限。比較問的是它與基準逐個案例有何不同。把 ../support-change/compare.yaml 複製到專案中;它包含一條比較規則:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: label-no-regression
    kind: non_inferiority
    metric: exact_label
    margin: 0.10
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:de76... of run_01M4...TJAD against run_01M4...ECVEF · 18 paired cases
exact_label: +22.2 points [-12.9, +57.0] · 18 paired · 0 missing · 0 excluded
format_valid: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
pii_free: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
Decisions
  label-no-regression  exact_label  non-inferiority, margin 10.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 3 more paired cases would decide it, if the difference holds (21 in total at 22% discordance)
Gate: BLOCK (exit 3)

候選修復了四個案例,沒有弄壞任何一個,估計提升 22 個百分點。但只有四個案例發生了變化,18 個配對案例留下的區間從差 12.9 個百分點到好 57 個百分點,跨越了 10 個百分點的邊際。這次比較尚不能排除候選比你所能接受的程度更差,所以結果是 INSUFFICIENT_EVIDENCE,以 3 結束。它下面那一行是樣本數估算。不帶 --policy 時,compare 會印出差異,說明專案的 release.yaml 沒有宣告任何比較規則,並以 0 結束,因為什麼都沒有被決策。

將候選與基準進行比較解釋了邊際以及其他規則類型。

故障排除

症狀原因與修復
app 的 ModuleNotFoundError在包含 app.py 的目錄中執行,或給 callable 一個可以從那裡導入的模組路徑。
規則指定了一個沒有任何評估器產出的指標規則的 metric 必須等於某個評估器的 criterion;報錯會列出現有的指標。
你編輯了分類器,但執行復用了所有輸出快取跟隨可呼叫物件的原始碼;它讀取的輔助檔案必須列在 system.code_paths 下。
exact_label 報告有案例缺失這些執行拋出了異常或超時;oloproof inspect RUN_ID --failures 會顯示每個錯誤。
估計值很高,執行卻以 3 結束起決定作用的是區間,而不是估計值。增加案例,或接受一個更低的下限,並在執行之前決定。

侷限

  • SDK 和 YAML 按評估器報告通過率。對於像這樣的分類器,沒有混淆矩陣,也沒有按類別的精確率和召回率;對於打分模型,predictive: 塊會提供這些(預測模型)。
  • observed_count 規則描述的是你執行的案例;它們不對未見過的輸入作任何斷言。
  • 基於 18 個案例的比較只能分辨很大的差異。五十個或更多真實案例是更有用的下限。
  • oloproof.yaml 無法指定自定義的 @evaluator;那需要 SDK(SDK)。