指南
教學:評估一個 RAG 應用
一份針對檢索增強應用的可執行演練,分兩條路徑:一個已有的黑盒應用,你從外部記錄它的檢索、上下文和引用;以及一個分階段的應用,Oloproof 逐階段執行它,從而讓 oloproof diagnose 能在受控變更下重新執行失敗的案例。兩者都在本機執行,無需提供者憑據。
每一步背後的概念(階段、相關性標籤、標準上下文、四種失敗標籤)在 RAG 評估頁面上;案例、評估器、指標、區間和閘門這些術語在核心概念中。本頁是貫穿它們的動手路線。
哪條路徑適合你
| 你的應用 | 路徑 | 你能得到什麼 | 你得不到什麼 |
|---|---|---|---|
| 一次呼叫進,一個回答出(一個服務、一個 HTTP 端點、一條你不想拆開的框架鏈) | A,黑盒 | 檢索指標、引用檢查、依據性評審、閘門、比較 | 受控干預:diagnose 不重新執行任何東西 |
| 可以分別呼叫的檢索和生成 | B,分階段 | A 中的一切、按階段快取,以及帶標準上下文、top-k 和重排序器並與對照組並列的 diagnose | 這三種之外的干預 |
如果不確定,就從 A 開始。它不需要修改應用,以後轉到 B 時,資料集、評估器和政策都可以保留。
前提條件
- Python 3.11 或更高版本,並已安裝 Oloproof(pip install oloproof)。
- 範例專案,它們隨軟體包一起提供:路徑 A 用 blackbox_rag,路徑 B 用 support_rag。把其中一個複製到新目錄並在那裡工作:
oloproof init --example blackbox_rag my-rag
cd my-rag下面的每條命令都在複製出的目錄中執行。執行、評判結果和診斷都儲存在那裡的 .oloproof/ 中。
路徑 A:把已有應用作為黑盒
檔案
| 檔案 | 它是什麼 |
|---|---|
| app.py | support_api(question),代替你的應用;以及 run(case),即適配器 |
| server.py | 通過 HTTP 提供的同一個應用,用於下面的 HTTP 變體 |
| data/corpus.jsonl | 應用所搜索的、包含 14 個段落的知識庫 |
| data/support.jsonl | 15 個案例:13 個帶有相關性標籤和標準段落,2 個兩者都沒有 |
| oloproof.yaml | 套件:資料集、系統、評估器、切片 |
| oloproof.http.yaml | 針對 HTTP 伺服器的同一個套件 |
| release.yaml | 單次執行的發布政策 |
| compare.yaml | 把候選執行與基準進行比較的政策 |
應用返回什麼
support_api 的行為就像一個你已經擁有的應用:它進行搜索,從符合字數預算的最佳來源建置提示,作答並引用。它的響應已經帶有它所做的事:
{
"answer": "Team plans include five seats.",
"cited": ["kb-03"],
"sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
"prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
"skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}你的應用的欄位名會有所不同。重要的是,它能針對每個問題告訴你它檢索到的排序後的來源、到達模型的來源,以及它引用的來源。如果做不到,先把這些加入它的響應或日誌:Oloproof 量測的是被記錄下來的內容,絕不會從回答中推斷檢索。
適配器
run 原樣呼叫應用,並把響應映射為三種有類型的產物,也就是檢索和引用評估器所讀取的記錄:
@system(
name="support-rag-blackbox",
version="tutorial",
records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
response = support_api(str(case["question"]))
recorder = current_case()
recorder.retrieval(
Retrieval(
query=case["question"],
depth=SEARCH_DEPTH,
candidates=tuple(
Passage(doc_id=s["id"], score=s["score"], text=s["text"])
for s in response["sources"]
),
)
)
recorder.context(
Context(
items=tuple(
ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
for i in response["prompt_sources"]
),
dropped=tuple(
DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
for i in response["skipped"]
),
token_budget=PROMPT_WORD_BUDGET,
)
)
recorder.citations(response["cited"])
return {"answer": response["answer"], "citations": response["cited"]}| 產物 | 結構 | 讀取者 |
|---|---|---|
| retrieval/v1 | query、depth,以及按檢索器返回順序排列的 candidates,每個都是一個 Passage(doc_id, chunk_id, score, text) | hit_rate、recall、mrr、ndcg |
| context/v1 | 到達模型的 items(doc_id、position、tokens、text),帶有 top_k 或 token_budget 這一 reason 的 dropped 項,以及 token_budget | citation_validity、groundedness_judge、citation_support_judge |
| citations/v1 | ids,每個是一個 doc_id 或 doc_id#chunk_id | citation_validity、citation_support_judge |
Oloproof 按給定的位置記錄,絕不重新排序。格式錯誤的產物會讓執行以結束代碼 2 停止,而不會被儲存。case 是案例的 input 物件,所以 case["question"] 就是資料集中的問題。
要使用你自己的應用,把 support_api 的函式體替換為對它的呼叫(一次 SDK 呼叫、一個 HTTP 請求),並保留 run。在 oloproof.yaml 中把 system.callable 以 module:function 的形式指向它。
HTTP 變體
HTTP 系統無法呼叫記錄器,所以改由它的響應攜帶證據,並且已經是上面的三種結構,設定則指明證據所在的位置:
system:
name: support-rag-http
version: tutorial
http:
url: http://127.0.0.1:8766/answer
output_path: result
artifacts:
retrieval/v1: evidence.retrieval
context/v1: evidence.context
citations/v1: evidence.citationsserver.py 提供的正是這些。啟動它,然後針對它執行:
python server.py 8766
oloproof run --config oloproof.http.yaml案例輸入作為 JSON 請求體發送。output_path 從響應中取出輸出,每個 artifacts 條目把一個點分路徑記錄為該類型;缺失或格式錯誤的欄位會讓執行以結束代碼 2 停止。結果與下面的可呼叫物件路徑完全相同。在你自己的服務中,證據物件通常是一個你為評估流量啟用的調試欄位。
案例宣告瞭什麼
{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}- expected.relevant 列出回答該問題的段落。檢索指標讀取它。沒有它的案例,比如 office_hours,會以 no_relevance_labels 被排除在這些指標之外:它離開分母,而不是被計為通過或失敗。
- expected.gold_context 是段落文字本身。路徑 A 從不使用它;路徑 B 在診斷時用它替換檢索到的上下文。
在實踐中,未標註的案例很常見,因為標註相關性需要工作量。它們仍然計入回答和引用檢查。
選擇評估器
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: hit_rate, k: 2}
- {type: recall, k: 2}
- {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4- contains 檢查回答是否包含預期文字。它是任務檢查:使用者是否得到了正確的答案。措辭多變時,改用精確匹配或評分準則評審。
- k: 2 下的 hit_rate 和 recall 在應用實際放進提示的深度上量測檢索。在模型從未看到的深度上的檢索指標,描述的是索引,而不是應用。
- citation_validity 檢查每個被引用的 id 是否指向一個到達了模型的段落;require_citations: true 還會讓什麼都不引用的回答失敗。
- groundedness_judge 和 citation_support_judge(可選)詢問模型回答是否得到上下文的支援。它們需要一個提供者、一個模型和環境變數中的憑據,並且每個案例都要花錢;評審在獲准做閘門之前必須達到什麼要求,見評審。
這裡無法使用 relevant_position 和 context_truncated 切片:它們把位置與應用的 top-k 進行比較,而只有分階段的系統才會宣告 top-k。請求它們會讓執行停止,並顯示 slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system。
發布政策
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answer-floor
metric: answer_correct
min: 0.70
- id: retrieval-floor
metric: hit_rate_at_2
min: 0.80
- id: citations-valid
metric: citations_valid
kind: observed_count
max_failures: 0min 規則只有在整個區間都越過下限時才通過,在整個區間都低於下限時失敗,否則為 INSUFFICIENT_EVIDENCE。observed_count 規則依據實際執行的案例作出決策,沒有區間:“這個套件中沒有無效引用”。見閘門。
執行
oloproof runRun run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 73.3% │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3% │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 miss如何閱讀:
- Gate: BLOCK (exit 1):有一條規則 FAIL 了。結束代碼 1 表示出現 FAIL;結束代碼 3 表示閘門在沒有 FAIL 的情況下阻止了發布(這裡會是 INSUFFICIENT_EVIDENCE);結束代碼 0 表示沒有出現任何政策要阻止的情況。[DECIDED/COMPLETE] 是執行狀態:每個案例都執行了。
- citations-valid FAIL:有一個回答什麼都沒有引用,而 require_citations 把這算作無效。
- answer-floor 是 INSUFFICIENT_EVIDENCE 而不是 PASS,儘管 73.3% 高於 70%:只有 15 個案例時,區間向下延伸到 44.8%,所以證據無法證明達到了下限。
- hit_rate_at_2 顯示 2 excluded:就是那兩個未標註的案例。它的分母是 13,而不是 15。
- 後面的 Slices 表是探索性的,從不用於閘門;低於 min_slice_support 的切片不顯示區間。
檢視失敗項
執行 id 在執行輸出的第一行。
oloproof inspect RUN_ID --failures4 of 15 cases failed, errored or did not finish
refund_review
output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
answer_correct: failed
money_back
output: {"answer": "I could not find that in the knowledge base.", "citations": []}
answer_correct: failed
hit_rate_at_2: failed
recall_at_2: failed
citations_valid: failed
security_review
output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
answer_correct: failed
seat_count
output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
answer_correct: failedoloproof inspect RUN_ID --case refund_review 印出一個案例的輸入、預期值、輸出和每個評判結果。記錄下來的產物在導出的包中:
oloproof export RUN_ID.oloproof/bundles/RUN_ID/cases.jsonl 的每一行是一個案例的記錄;它的 artifacts 欄位保存了被記錄的內容。對於 money_back,該欄位內容為:
{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}僅憑記錄下來的證據來解讀這四個失敗:
| 案例 | 記錄顯示了什麼 | 有意義的下一步 |
|---|---|---|
| money_back | 檢索什麼都沒有返回:問題與退款段落沒有任何共同的詞 | 查詢改寫或同義詞,用 hit_rate_at_2 來量測 |
| refund_review、security_review | hit_rate_at_2 通過了,但回答來自另一個段落 | 檢視 context/v1:相關段落是否因為預算而被丟棄了? |
| seat_count | 正確的段落被檢索到、被保留並被引用;回答說“five”,案例期望的是“5” | 修復期望值或回答格式,而不是檢索 |
這張表是你對記錄的解讀。它是失敗與某個階段之間的關聯,而不是已經證明的原因:沒有任何東西在改變該階段後重新執行過這個案例。
diagnose 對黑盒做了什麼
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc每個案例都是 UNRESOLVED,原因為 intervention_unsupported。Oloproof 無法把標準段落交給黑盒來代替它自己的檢索,所以它不會假裝能做到。受控干預需要路徑 B。
做一次候選變更並比較
記錄表明 money_back 在檢索階段失敗。候選變更會在搜索之前用同義詞擴充問題。在 app.py 中:
EXPAND_QUERY = True修改代碼會改變該執行所記錄的系統版本。再執行一次,然後把候選與基準進行比較:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml單看候選執行:citations-valid 現在 PASS,hit_rate_at_2 顯示 100.0% [75.2%, 100.0%],而閘門仍以結束代碼 3 阻止,因為 answer-floor 和 retrieval-floor 仍是 INSUFFICIENT_EVIDENCE。比較結果:
Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)這次變更修復了它針對的案例(多答對一個,在 15 個配對案例上 +6.7 個百分點)。比較仍然無法確定候選比基準差的程度不超過邊際:15 個配對案例留下的區間寬約 67 個百分點。規劃行說明,如果差異保持不變,還需要多少配對案例才能作出決策。下一步是更大的套件,而不是換一個邊際。見比較兩次執行和比較規則。
路徑 B:帶診斷的分階段應用
分階段的檔案
路徑 B 執行 support_rag 範例,它在 RAG 評估頁面上有介紹。複製它:
oloproof init --example support_rag my-staged-rag
cd my-staged-rag| 檔案 | 它是什麼 |
|---|---|
| app.py | SupportRag,一個用 @rag_system 裝飾的類:retrieve(input, depth)、generate(input, context)、count_tokens(passage) |
| data/corpus.jsonl、data/support.jsonl | 知識庫,以及 13 個案例,每個都帶有 relevant 和 gold_context |
| oloproof.yaml | system.rag 指向這個類,並設置 depth、top_k、token_budget、index_version |
| release.yaml、compare.yaml | 與路徑 A 相同的政策 |
與路徑 A 的區別在於由誰來組裝上下文。在這裡,Oloproof 呼叫 retrieve,保留前 top_k 個候選,丟棄超出 token_budget 的段落,並把其餘的傳給 generate。由於它把各階段分開,它可以分別快取它們,並用不同的上下文重新執行生成。要適配你自己的應用,替換 retrieve(呼叫你的索引,按檢索器的順序返回 Retrieval(candidates=[Passage(...)]))和 generate(用給定的段落呼叫你的模型)的函式體。把 index_version 設為一個會隨索引變化而變化的值:它是檢索身份的一部分,過時的值會讓已快取的檢索結果被複用,而當前索引已經不再返回它們了。
同樣的設定還允許使用 relevant_position 和 context_truncated 切片,以及一個覆蓋完整檢索深度的 ndcg 評估器。
執行分階段的套件
oloproof runGate: BLOCK (exit 3)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ PASS │ observed_failures_within_limit │
│ answer_correct │ 69.2% │ [38.5%, 91.0%] │ 9 / 13 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ ndcg_at_6 │ 0.866 │ [0.506, 0.990] │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0% │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 missStages 行是分階段系統自己的快取。結束代碼 3:沒有任何規則 FAIL,但有兩條規則缺少 PASS 所需的證據。
用標準上下文進行診斷,並與對照組並列
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…從失敗的案例中會生成兩個子執行:一個用案例的 gold_context 代替檢索到的上下文,另一個是原樣重新執行它們的對照組。對照組讓解讀變得可靠:一個在普通重跑中就通過的案例是不穩定的,而不是被診斷出來的。無論發現什麼,diagnose 都以 0 結束;它不對發布作任何決策。
oloproof inspect DIAGNOSIS_IDmoney_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2解讀標籤
| 標籤 | 觀測到了什麼 | 它不能確立什麼 |
|---|---|---|
| RETRIEVAL_MISS | 沒有檢索到相關段落,而案例在使用標準段落時通過了 | 檢索是唯一的問題,或者某個特定的檢索變更就能修復它 |
| RANKED_OUT | 檢索到了一個相關段落,但排在 top_k 之後,而案例在使用標準段落時通過了 | 擴大 top-k 對其他案例也會有幫助 |
| CONTEXT_ASSEMBLY_LOSS | top-k 之內的一個相關段落從上下文中被丟棄了,而案例在使用標準段落時通過了 | 多大的預算才夠 |
| GENERATION_FAILURE | 即使手握標準段落,案例仍然失敗 | 出錯的是模型,而不是提示或期望值 |
| UNRESOLVED | 無法得出任何結論:系統不是分階段的(intervention_unsupported),案例在對照組下恢復了(unstable_under_control),它沒有相關性標籤(no_relevance_labels),或證據缺失 | 關於該案例的任何事情 |
每個標籤都是在對這些案例施加一種干預的情況下,失敗與某個階段之間的關聯。它不是已經證明的原因:“Implicated”和“Candidate experiment”是輸出中使用的最強的措辭,而計數只描述所選的案例(“no population claim”)。seat_count 是一個很好的提醒:它在拿到正確段落時仍然失敗,因為知識庫寫的是“five”而案例期望的是“5”,這是任何檢索變更都無法修復的。
有標準段落和沒有標準段落的案例
只有宣告瞭 expected.gold_context 的失敗案例才能被重新執行。從 seat_count 和 money_back 中移除標準段落(並移除 money_back 的相關性標籤),同一條命令會報告:
Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the context被排除的案例會被列出,而不是被悄悄丟棄。另外請注意,移除相關性標籤對執行本身有什麼影響:hit_rate_at_2 升到了 100.0%(12 / 12 observed,1 excluded),因為檢索遺漏的那一個案例不再被量測。未標註的案例離開分母;它們不計為通過,而基於更少案例的指標可能看起來比應用的實際情況更好。先標註難的案例。
在動手之前先檢驗修復:top-k 和重排序器
另外兩種干預會用不同的設置重放記錄下來的檢索,所以不會再次呼叫檢索器:
oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correctRecovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recovered重排序器是你編寫的一個函式 (input, candidates) -> candidates。把它保存為 app.py 旁邊的 rerank.py:
"""A candidate reranker: shorter passages first, so more of them fit the token budget."""
from oloproof import Passage
def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
return sorted(candidates, key=lambda passage: len((passage.text or "").split()))oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correctRecovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recovered兩者都沒有恢復任何案例,這正是標準上下文標籤所預測的:這裡沒有哪個失敗是因為段落剛好排在截斷線之下。每次重放都會延續標準上下文標籤,所以這些診斷可以放在一起閱讀。
執行診斷指出的實驗,並進行比較
在 oloproof.yaml 中把 token_budget 提高到 120,然後:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlStages: retrieve 13 hit/0 miss; generate 7 hit/6 miss每次檢索都被複用了,因為 top_k 和預算不屬於檢索的身份;只有上下文發生了變化的六個案例被重新生成。
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
Gate: BLOCK (exit 3)這次實驗沒有幫助:沒有一個案例改變了結論,所以診斷提出的假設在這些案例上沒有得到支援。這是一個有用的結果。下一個實驗是更小的分塊,或者那兩個上下文組裝案例的提示;seat_count 需要修復它的期望值。
故障排除
| 症狀 | 原因 | 修復 |
|---|---|---|
| Configuration error: slice 'relevant_position' ... needs a staged system | 在可呼叫物件或 HTTP 系統上使用了位置切片 | 去掉這個切片,或轉到路徑 B |
| 執行以結束代碼 2 停止,並顯示 malformed retrieval/v1 artifact | 出現了模式不允許的欄位,或候選數量超過了 depth | 只映射文件中列出的欄位;把 depth 設為至少等於返回的數量 |
| citations_valid 顯示 0 / 0 observed · 15 missing,它的規則是 INSUFFICIENT_EVIDENCE,原因為 no_observations | 適配器沒有記錄 citations/v1(或 context/v1);每個這樣的案例都是缺失,而不是通過 | 在適配器的每條路徑上都記錄兩者,包括“沒有回答”;oloproof inspect RUN_ID --failures 會顯示每個案例的錯誤 |
| 某個檢索指標顯示很多 excluded | 案例沒有 expected.relevant | 標註它們,或者在知情的情況下接受更小的分母 |
| diagnose 顯示 UNRESOLVED ... not staged | 路徑 A | 符合預期;要做干預請使用路徑 B |
| diagnose 拒絕執行並顯示 an intervention must re-execute the same system | 自該執行以來代碼或設定發生了變化 | 診斷當前版本的執行,或恢復執行時的那個版本 |
| 診斷選中的案例比失敗的少 | 失敗的案例沒有 expected.gold_context | 添加標準段落;被排除的案例會在輸出中列出 |
| 索引變化後檢索結果仍被複用 | index_version 沒有變化 | 索引變化時修改 index_version |
侷限
- Oloproof 呼叫你的應用;它不託管、不隔離也不重置它。它的索引、快取以及它保存的任何狀態都歸你管理。
- 在黑盒上無法進行干預:diagnose 把每個案例都標為 UNRESOLVED,並且不重新執行任何東西。
- 干預包括標準上下文、top-k 和重排序器。沒有分塊、嵌入或提示方面的干預。
- 診斷標籤描述的是在一種干預下、與對照組並列的所選失敗案例。它們把失敗與某個階段關聯起來;它們不證明原因,也不對未被選中的案例作任何斷言。
- 檢索指標需要相關性標籤,診斷需要標準段落;Oloproof 不會建立這兩者。
- 確定性範例代替了真實的檢索器和模型。在 generate 中使用真實模型或使用評判評估器,都會呼叫提供者,需要憑據,並且每個案例都要花錢。
- 哪些功能在哪裡可用(SDK、YAML 與瀏覽器)見目前可用的功能。