跳至主要內容

指南

教學:使用評分準則評審的文字生成

評估一個編寫自由文字的函式,這裡是一個工單摘要器,使用格式檢查和一個評分準則評審;在評審獲准決定任何事情之前,先用一個人的標籤來量測它;然後比較一次真正的修改。評審在本機上執行,無需模型,也無需聯網,另有一個可選步驟會換用真實模型。

你將建置什麼

一個把客服工單轉寫成一兩句話的摘要器。“好”是一個判斷,而不是字串匹配,所以任務成功由一個帶評分準則的 LLM 評審來決定:摘要是否說明了客服人員需要的事實?兩個確定性評估器檢查格式,這不需要參考。案例、執行、指標、評審和閘門等術語在核心概念中有定義。

同樣的結構也適用於資訊抽取或任何其他生成:一個函式在字典中返回文字,參考說明好的回答必須包含什麼,評分準則說明如何判定。

前提條件

  • Python 3.11 或更高版本,以及安裝在虛擬環境中的 Oloproof:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • 範例專案,它隨軟體包一起提供。把它複製到一個新目錄並在那裡工作:
oloproof init --example generation ticket-summaries
cd ticket-summaries
  • 為替身評審空出 8799 端口(如果不行,就在兩處同時修改)。

在“可選:用真實模型作為評審”之前的每一步都是離線且確定性的:沒有 API 密鑰,沒有提供者帳號,沒有費用。

檔案

ticket-summaries/
  app.py                        the summariser under test (baseline)
  app_v2.py                     the candidate change
  judge_server.py               a stand-in judge speaking the OpenAI API on 127.0.0.1
  rubrics/covers_facts.md       the judge's rubric
  oloproof.yaml                 the suite
  release.yaml                  rules for a run
  compare.yaml                  a rule for a comparison
  data/tickets.jsonl            20 cases
  labels/reviewer_verdicts.csv  one person's verdicts on the baseline's summaries
  fill_labels.py                copies those verdicts into a labelling sheet

在 ticket-summaries/ 中執行每條命令。

替身評審,以及它不是什麼

評分準則評審是一種評估器,它把一個提示(評分準則、案例的輸入、它的 expected 和輸出)發送給一個模型,並讀回 {"pass": true|false, "rationale": "..."}。Oloproof 可以與任何使用 OpenAI chat API 的伺服器通信,而 localhost 上的伺服器不需要密鑰。

judge_server.py 就是這樣一個伺服器,但它不是模型。只有當摘要包含案例 expected 中 must_mention 下的每一個短語(忽略大小寫)時,它才讓摘要通過。這是一條固定規則,所以本教學在每臺機器上都給出相同的數字。它無法察覺編造的事實,而真實的模型評審會被要求這樣做。在第二個終端機中啟動它並讓它保持執行:

python judge_server.py --port 8799
stand-in judge on http://127.0.0.1:8799/v1

應用及其適配器

# app.py
@system(name="ticket-summariser", version="first-sentence")
def summarise(case: dict[str, Any]) -> dict[str, str]:
    return {"summary": sentences(str(case["ticket"]))[0]}

Python 應用的適配器就是這個函式:它接收案例的 input 並返回一個字典。對於你自己的生成器,在函式內部呼叫你的模型或鏈,並把文字放在某個鍵下返回。Oloproof 對每個案例呼叫它一次,並按函式的原始碼和宣告的 version 快取輸出;它不管理你的模型客戶端、提示詞或狀態。把函式讀取的檔案(例如提示模板)列在 system.code_paths 下。

資料集

{"id":"t01","input":{"ticket":"Hello. Order 1042 arrived with a cracked screen. I would like a replacement, not a refund."},"expected":{"must_mention":["1042","cracked","replacement"]}}
{"id":"t06","input":{"ticket":"Please cancel my subscription at the end of this month. I am moving abroad."},"expected":{"must_mention":["cancel","end of this month"]}}

input 是函式接收到的內容。expected 是評審讀取的參考:這裡是摘要必須包含的事實清單,而不是一份完整的參考摘要,因為許多不同的摘要都是正確的。t01 的輸出是 {"summary": "Hello."}。

選擇評估器

version: 1
project: ticket-summaries
dataset: data/tickets.jsonl
system:
  name: ticket-summariser
  version: first-sentence
  callable: app:summarise
  timeout_s: 30
evaluators:
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [summary]
      properties:
        summary: {type: string, minLength: 1}
      additionalProperties: false
  - type: regex
    criterion: short_enough
    field: summary
    pattern: '^.{1,160}$'
    pass_if: match
  - type: rubric_judge
    criterion: covers_facts
    provider: openai_compatible
    model: stand-in-judge
    base_url: http://127.0.0.1:8799/v1
    rubric_file: rubrics/covers_facts.md
判據評估器需要 expected量測
format_validjson_schema否格式:一個非空字串欄位
short_enoughregex否格式:最多 160 個字符
covers_factsrubric_judge是任務成功,按評分準則的定義

Hello. 通過了兩項格式檢查。只有評審會說它是一個無用的摘要。評審也可以在沒有參考的情況下執行:像“PASS if the summary contains no greeting”這樣的評分準則只讀取輸入和輸出,沒有 expected 的案例仍然會被評判。這時它做不到的,是對照你信任的答案核對事實。

評分準則:

PASS when the summary states every fact listed under must_mention in the expected answer, in
words a support agent would recognise, and adds nothing the ticket does not say.
FAIL when any listed fact is missing, changed or contradicted.

政策

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: short-enough
    metric: short_enough
    kind: observed_count
    max_failures: 0
  - id: covers-facts-floor
    metric: covers_facts
    min: 0.60

require_validated_evaluators: true 是引擎的預設值,這裡把它寫出來,是因為它正是本教學的要點:一個沒有人與人工比較過的評審,不能決定一條規則。

執行

oloproof run
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format       │ format_valid │ PASS                  │ observed_failures_within_limit │
│ short-enough       │ short_enough │ PASS                  │ observed_failures_within_limit │
│ covers-facts-floor │ covers_facts │ INSUFFICIENT_EVIDENCE │ evaluator_not_validated        │
covers-facts-floor: the judge (or model or custom evaluator) behind this rule has not been measured against
people yet, so it may not decide.
  Label a sample:  oloproof review run_01M4... --criterion covers_facts --by YOU --sample 20
  Then measure it: oloproof evaluators validate EVALUATOR_ID --by YOU (ids: oloproof evaluators list)
│ format_valid │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ short_enough │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ covers_facts │ 45.0%    │ [23.0%, 68.5%]  │ 9 / 20 observed · 0 missing · 0 excluded  │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 miss

格式規則通過了。評審讓 20 個摘要中的 9 個通過,但規則是 INSUFFICIENT_EVIDENCE,原因為 evaluator_not_validated,閘門以結束代碼 3 阻止。規則並沒有依據 45% 作出決策:評審的錯誤率在被量測之前是未知的,所以基於它的結論建置的區間會帶有一個未說明的誤差。引擎把這報告為 INSUFFICIENT_EVIDENCE,而不是 MANUAL_REVIEW 或 FAIL:缺少的是用來決策的證據,輸出會印出出提供證據的兩條命令。

檢視失敗項

oloproof inspect RUN_ID --failures
11 of 20 cases failed, errored or did not finish

t01
  output: {"summary": "Hello."}
  covers_facts: failed
    judge text, not verified: missing: 1042, cracked, replacement

t02
  output: {"summary": "I was charged twice for order 2210."}
  covers_facts: failed
    judge text, not verified: missing: 49
...

評審的理由顯示為“judge text, not verified”:它是模型的解釋,而不是證據。無論如何,規律很清楚:第一句往往是一句問候。

用一個人來量測評審

驗證會把評審的結論與一個人對同一批迴答的結論進行比較。從執行的案例中隨機抽取一個樣本放進一張表。評審的結論不會出現在表中,這樣標註者就不會被它們錨定:

oloproof labels export RUN_ID --criterion covers_facts --sample 20 --local --out sample.csv
Wrote 20 cases to sample.csv, drawn at random with seed 2701013296, without the judge's verdict.
  This is a local sample, good-faith only, because it was drawn on this machine.
Fill in `passed` (pass or fail) and `labelled_by` on each row you judge, then run `oloproof labels import sample.csv`.

--local 在本機上抽樣,而不詢問託管工作區;種子由引擎選擇。只有 20 個案例時,20 個的樣本就是全部案例。實際操作中,一個人會讀每一行的工單和摘要並填寫 passed。在本教學中,labels/reviewer_verdicts.csv 保存了一位審核者對基準摘要給出的結論,fill_labels.py 把它們複製到表中:

python fill_labels.py sample.csv
oloproof labels import sample.csv
filled 20 rows of sample.csv
Recorded 20 labels from sample.csv (20 measurement).

審核者與評審有一次意見不同:在 t02(“I was charged twice for order 2210.”)上,他們認為缺少金額無關緊要,讓它通過了。標籤指明瞭它們所評判的確切回答,所以這些結論只適用於基準執行。

找到評審的版本 id 並驗證它:

oloproof evaluators list
oloproof evaluators validate EVALUATOR_ID --by alice
covers_facts  LLM_JUDGE  UNVALIDATED  (declared)  sha256:a662...

covers_facts: sha256:a662... is now VALIDATED
  agreement 95.0% [75.1%, 99.9%] · 19 of 20 labelled cases agreed · 0 labelled but not judged · kappa 0.900
  bias -5.0 points [-32.4, +20.7] · the judge's pass rate minus the people's · 20 cases · 0 labelled but not judged
  passes what people pass 90.0% [55.4%, 99.8%] · the judge passed 9 of 10 cases people passed · 0 labelled but not judged
  fails what people fail 100.0% [69.1%, 100.0%] · the judge failed 10 of 10 cases people failed · 0 labelled but not judged

讀區間,而不是 95%:20 個標籤顯示一致率至少為 75.1%。政策可以用 minimum_evaluator_agreement 提出更高的要求,它比較的正是這個下界,而 validate 會拒絕低於它的評審。評審指南介紹了這個標準、偏差、探針,以及在終端機中進行標註的 oloproof review。

現在不呼叫摘要器或評審,對已儲存的執行重新做出決策:

oloproof gate RUN_ID --policy release.yaml
valid-format: PASS (observed_failures_within_limit)
short-enough: PASS (observed_failures_within_limit)
covers-facts-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)
  no sample size would make this PASS: the observed rate (0.500) is itself below the threshold (0.600), so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

現在評審可以決策了,而決策針對的是摘要器:它引用的比率 0.500 並不是評審的 45%。由於這次執行有一個盲測的隨機量測標籤樣本,閘門讀取的是經這些標籤校正後的評審(見評審指南中的“Judge-corrected gates”)。這種校正是 PPI,即預測驅動推斷:它用已標註的樣本數測評審的比率與人工的比率相差多遠,並據此移動估計值、加寬區間。導出檔案中關於 PPI 的說明指的也是這一點。無論哪種方式,基準都沒有達到下限,增加案例也改變不了這一點。

做一次真正的修改

app_v2.py 會跳過簡短的客套話,保留接下來的兩句。把它複製到 app.py 上覆蓋,在 oloproof.yaml 的 system 下設置 version: skip-pleasantries,保持評審執行,然後:

oloproof run
Gate: ALLOW (exit 0)
│ covers-facts-floor │ covers_facts │ PASS  │ lower_bound_meets_minimum      │
│ covers_facts │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 6 hit/54 miss

評審是同一個已驗證的版本,所以它的規則直接作出決策。有六個評判結果來自快取,針對的是兩個版本寫得完全相同的摘要。沒有人標註過這些新摘要;是評審的驗證讓它的結論得以成立。

把候選與基準進行比較

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
  - id: covers-more-facts
    kind: superiority
    metric: covers_facts
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
format_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
short_enough: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
covers_facts: +55.0 points [+13.0, +84.4] · 20 paired · 0 missing · 0 excluded
Decisions
  covers-more-facts  covers_facts  superiority  PASS  difference_above_zero
Gate: ALLOW (exit 0)

比較不應用 PPI 校正:它比較的是評審自己在兩次執行上的結論,這就是為什麼提升是從評審的 45% 而不是上面校正後的 0.500 算起。十一個摘要變好了,沒有一個變差;提升的區間完全在零以上,所以優效性規則通過,命令以 0 結束。格式由不允許任何失敗的執行規則來守護,而不是由比較來守護:在 20 個案例上,比較兩個完美的格式分數,只能說差異在 23.6 個百分點以內。

可選:用真實模型作為評審

這一步離開了離線路徑。它需要一個模型伺服器,如果使用雲端提供者,還需要密鑰和費用。

  • 本機,無需密鑰也無需費用:localhost 上的 Ollama、LM Studio 或 llama.cpp。拉取一個聊天模型(對於 Ollama,ollama pull llama3.1)。
  • 雲端:provider: anthropic 或 openai,並用 api_key_env 指明保存你密鑰的變數;或者 openai_compatible,帶上 base_url 和 api_key_env。每個案例是一次評判呼叫(第一次回覆不是有效 JSON 時為兩次),按你的提供者的費率計費,並且對於已經評判過的回答,Oloproof 絕不會再次呼叫評審。

把草擬的評審寫在一個單獨的檔案中,就像它出現在 evaluators: 下那樣:

# live_judge.yaml
type: rubric_judge
criterion: covers_facts
provider: openai_compatible
model: llama3.1
base_url: http://localhost:11434/v1
rubric_file: rubrics/covers_facts.md

並用審核者已經標註過的回答來試驗它,而不驗證也不採用它:

oloproof evaluators try live_judge.yaml

本機伺服器預設一次只應答一個請求;在 oloproof.yaml 中加入 concurrency: {system: 2, judge: 2},這樣排隊的呼叫就不會超時。在一臺筆記本電腦上用一個小型本機模型(qwen2.5vl)執行這一步,印出出:

covers_facts: draft sha256:b88a... on 20 labelled cases · 20 judged now, 0 from cache, 11 errored
  agreement 88.9% [19.1%, 99.9%] · 8 of 9 labelled cases agreed · 11 labelled but not judged · kappa 0.769

有十一次呼叫超時,一致性區間把每一次都按兩種方向計入,所以它向下延伸到 19.1%:不應答的評審是無法被量測的。更大的模型、更長的超時或更少的併發呼叫可以解決這個問題。要採用這個模型,就把它放進 oloproof.yaml,替換掉替身。那是一個新的評估器版本:它的設定(模型、端點、評分準則)就是它的身份,所以替身的驗證不會延續過來。用它重新執行基準,並像上面那樣對照標籤驗證它。

故障排除

症狀原因與修復
covers_facts 全部缺失,no_observations評判伺服器沒有執行,或不在 base_url 上。每次評判呼叫都出錯了;oloproof inspect RUN_ID --failures 會顯示原因。
驗證之後仍出現 evaluator_not_validated你修改了評審(模型、端點、端口、評分準則),產生了一個新版本。驗證那個版本。
labels import 拒絕檔案並指出某一行該行指定的案例或執行不在這次執行中;從你要標註的執行重新導出。
labels export 說無法連接到工作區你已登入某個工作區,所以它請求由工作區來抽樣。--local 改為在本機抽樣。
雲端評審在任何呼叫之前就失敗它的密鑰不在 api_key_env 指定的變數中。

侷限

  • 替身評審只是短語匹配。它演示的是工作流程,而不是評判品質。
  • 沒有 BLEU、ROUGE 或嵌入相似度評估器。在 SDK 中,可以用 @evaluator 寫一個;oloproof.yaml 目前還無法指定自定義評估器。
  • 評審看到的是文字:輸入、參考和輸出的 JSON。它看不到圖像或音頻。
  • 二十個標籤給出的一致性區間很寬。對於你所依賴的評審,請以隨機且盲測的方式標註更多。
  • 本機樣本只代表善意。對於其他人所依賴的評審,請推送執行,讓託管工作區來抽取樣本(評審)。