跳至主要內容

指南

教學:評估一個 RAG 應用

一份針對檢索增強應用的可執行演練,分兩條路徑:一個已有的黑盒應用,你從外部記錄它的檢索、上下文和引用;以及一個分階段的應用,Oloproof 逐階段執行它,從而讓 oloproof diagnose 能在受控變更下重新執行失敗的案例。兩者都在本機執行,無需提供者憑據。

每一步背後的概念(階段、相關性標籤、標準上下文、四種失敗標籤)在 RAG 評估頁面上;案例、評估器、指標、區間和閘門這些術語在核心概念中。本頁是貫穿它們的動手路線。

哪條路徑適合你

你的應用路徑你能得到什麼你得不到什麼
一次呼叫進,一個回答出(一個服務、一個 HTTP 端點、一條你不想拆開的框架鏈)A,黑盒檢索指標、引用檢查、依據性評審、閘門、比較受控干預:diagnose 不重新執行任何東西
可以分別呼叫的檢索和生成B,分階段A 中的一切、按階段快取,以及帶標準上下文、top-k 和重排序器並與對照組並列的 diagnose這三種之外的干預

如果不確定,就從 A 開始。它不需要修改應用,以後轉到 B 時,資料集、評估器和政策都可以保留。

前提條件

  • Python 3.11 或更高版本,並已安裝 Oloproof(pip install oloproof)。
  • 範例專案,它們隨軟體包一起提供:路徑 A 用 blackbox_rag,路徑 B 用 support_rag。把其中一個複製到新目錄並在那裡工作:
oloproof init --example blackbox_rag my-rag
cd my-rag

下面的每條命令都在複製出的目錄中執行。執行、評判結果和診斷都儲存在那裡的 .oloproof/ 中。

路徑 A:把已有應用作為黑盒

檔案

檔案它是什麼
app.pysupport_api(question),代替你的應用;以及 run(case),即適配器
server.py通過 HTTP 提供的同一個應用,用於下面的 HTTP 變體
data/corpus.jsonl應用所搜索的、包含 14 個段落的知識庫
data/support.jsonl15 個案例:13 個帶有相關性標籤和標準段落,2 個兩者都沒有
oloproof.yaml套件:資料集、系統、評估器、切片
oloproof.http.yaml針對 HTTP 伺服器的同一個套件
release.yaml單次執行的發布政策
compare.yaml把候選執行與基準進行比較的政策

應用返回什麼

support_api 的行為就像一個你已經擁有的應用:它進行搜索,從符合字數預算的最佳來源建置提示,作答並引用。它的響應已經帶有它所做的事:

{
  "answer": "Team plans include five seats.",
  "cited": ["kb-03"],
  "sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
  "prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
  "skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}

你的應用的欄位名會有所不同。重要的是,它能針對每個問題告訴你它檢索到的排序後的來源、到達模型的來源,以及它引用的來源。如果做不到,先把這些加入它的響應或日誌:Oloproof 量測的是被記錄下來的內容,絕不會從回答中推斷檢索。

適配器

run 原樣呼叫應用,並把響應映射為三種有類型的產物,也就是檢索和引用評估器所讀取的記錄:

@system(
    name="support-rag-blackbox",
    version="tutorial",
    records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
    response = support_api(str(case["question"]))
    recorder = current_case()
    recorder.retrieval(
        Retrieval(
            query=case["question"],
            depth=SEARCH_DEPTH,
            candidates=tuple(
                Passage(doc_id=s["id"], score=s["score"], text=s["text"])
                for s in response["sources"]
            ),
        )
    )
    recorder.context(
        Context(
            items=tuple(
                ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
                for i in response["prompt_sources"]
            ),
            dropped=tuple(
                DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
                for i in response["skipped"]
            ),
            token_budget=PROMPT_WORD_BUDGET,
        )
    )
    recorder.citations(response["cited"])
    return {"answer": response["answer"], "citations": response["cited"]}
產物結構讀取者
retrieval/v1query、depth,以及按檢索器返回順序排列的 candidates,每個都是一個 Passage(doc_id, chunk_id, score, text)hit_rate、recall、mrr、ndcg
context/v1到達模型的 items(doc_id、position、tokens、text),帶有 top_k 或 token_budget 這一 reason 的 dropped 項,以及 token_budgetcitation_validity、groundedness_judge、citation_support_judge
citations/v1ids,每個是一個 doc_id 或 doc_id#chunk_idcitation_validity、citation_support_judge

Oloproof 按給定的位置記錄,絕不重新排序。格式錯誤的產物會讓執行以結束代碼 2 停止,而不會被儲存。case 是案例的 input 物件,所以 case["question"] 就是資料集中的問題。

要使用你自己的應用,把 support_api 的函式體替換為對它的呼叫(一次 SDK 呼叫、一個 HTTP 請求),並保留 run。在 oloproof.yaml 中把 system.callable 以 module:function 的形式指向它。

HTTP 變體

HTTP 系統無法呼叫記錄器,所以改由它的響應攜帶證據,並且已經是上面的三種結構,設定則指明證據所在的位置:

system:
  name: support-rag-http
  version: tutorial
  http:
    url: http://127.0.0.1:8766/answer
    output_path: result
    artifacts:
      retrieval/v1: evidence.retrieval
      context/v1: evidence.context
      citations/v1: evidence.citations

server.py 提供的正是這些。啟動它,然後針對它執行:

python server.py 8766
oloproof run --config oloproof.http.yaml

案例輸入作為 JSON 請求體發送。output_path 從響應中取出輸出,每個 artifacts 條目把一個點分路徑記錄為該類型;缺失或格式錯誤的欄位會讓執行以結束代碼 2 停止。結果與下面的可呼叫物件路徑完全相同。在你自己的服務中,證據物件通常是一個你為評估流量啟用的調試欄位。

案例宣告瞭什麼

{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}
  • expected.relevant 列出回答該問題的段落。檢索指標讀取它。沒有它的案例,比如 office_hours,會以 no_relevance_labels 被排除在這些指標之外:它離開分母,而不是被計為通過或失敗。
  • expected.gold_context 是段落文字本身。路徑 A 從不使用它;路徑 B 在診斷時用它替換檢索到的上下文。

在實踐中,未標註的案例很常見,因為標註相關性需要工作量。它們仍然計入回答和引用檢查。

選擇評估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4
  • contains 檢查回答是否包含預期文字。它是任務檢查:使用者是否得到了正確的答案。措辭多變時,改用精確匹配或評分準則評審。
  • k: 2 下的 hit_rate 和 recall 在應用實際放進提示的深度上量測檢索。在模型從未看到的深度上的檢索指標,描述的是索引,而不是應用。
  • citation_validity 檢查每個被引用的 id 是否指向一個到達了模型的段落;require_citations: true 還會讓什麼都不引用的回答失敗。
  • groundedness_judge 和 citation_support_judge(可選)詢問模型回答是否得到上下文的支援。它們需要一個提供者、一個模型和環境變數中的憑據,並且每個案例都要花錢;評審在獲准做閘門之前必須達到什麼要求,見評審。

這裡無法使用 relevant_position 和 context_truncated 切片:它們把位置與應用的 top-k 進行比較,而只有分階段的系統才會宣告 top-k。請求它們會讓執行停止,並顯示 slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system。

發布政策

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: retrieval-floor
    metric: hit_rate_at_2
    min: 0.80
  - id: citations-valid
    metric: citations_valid
    kind: observed_count
    max_failures: 0

min 規則只有在整個區間都越過下限時才通過,在整個區間都低於下限時失敗,否則為 INSUFFICIENT_EVIDENCE。observed_count 規則依據實際執行的案例作出決策,沒有區間:“這個套件中沒有無效引用”。見閘門。

執行

oloproof run
Run run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct  │ 73.3%    │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3%    │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 miss

如何閱讀:

  • Gate: BLOCK (exit 1):有一條規則 FAIL 了。結束代碼 1 表示出現 FAIL;結束代碼 3 表示閘門在沒有 FAIL 的情況下阻止了發布(這裡會是 INSUFFICIENT_EVIDENCE);結束代碼 0 表示沒有出現任何政策要阻止的情況。[DECIDED/COMPLETE] 是執行狀態:每個案例都執行了。
  • citations-valid FAIL:有一個回答什麼都沒有引用,而 require_citations 把這算作無效。
  • answer-floor 是 INSUFFICIENT_EVIDENCE 而不是 PASS,儘管 73.3% 高於 70%:只有 15 個案例時,區間向下延伸到 44.8%,所以證據無法證明達到了下限。
  • hit_rate_at_2 顯示 2 excluded:就是那兩個未標註的案例。它的分母是 13,而不是 15。
  • 後面的 Slices 表是探索性的,從不用於閘門;低於 min_slice_support 的切片不顯示區間。

檢視失敗項

執行 id 在執行輸出的第一行。

oloproof inspect RUN_ID --failures
4 of 15 cases failed, errored or did not finish

refund_review
  output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
  answer_correct: failed

money_back
  output: {"answer": "I could not find that in the knowledge base.", "citations": []}
  answer_correct: failed
  hit_rate_at_2: failed
  recall_at_2: failed
  citations_valid: failed

security_review
  output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
  answer_correct: failed

seat_count
  output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
  answer_correct: failed

oloproof inspect RUN_ID --case refund_review 印出一個案例的輸入、預期值、輸出和每個評判結果。記錄下來的產物在導出的包中:

oloproof export RUN_ID

.oloproof/bundles/RUN_ID/cases.jsonl 的每一行是一個案例的記錄;它的 artifacts 欄位保存了被記錄的內容。對於 money_back,該欄位內容為:

{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}

僅憑記錄下來的證據來解讀這四個失敗:

案例記錄顯示了什麼有意義的下一步
money_back檢索什麼都沒有返回:問題與退款段落沒有任何共同的詞查詢改寫或同義詞,用 hit_rate_at_2 來量測
refund_review、security_reviewhit_rate_at_2 通過了,但回答來自另一個段落檢視 context/v1:相關段落是否因為預算而被丟棄了?
seat_count正確的段落被檢索到、被保留並被引用;回答說“five”,案例期望的是“5”修復期望值或回答格式,而不是檢索

這張表是你對記錄的解讀。它是失敗與某個階段之間的關聯,而不是已經證明的原因:沒有任何東西在改變該階段後重新執行過這個案例。

diagnose 對黑盒做了什麼

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc

每個案例都是 UNRESOLVED,原因為 intervention_unsupported。Oloproof 無法把標準段落交給黑盒來代替它自己的檢索,所以它不會假裝能做到。受控干預需要路徑 B。

做一次候選變更並比較

記錄表明 money_back 在檢索階段失敗。候選變更會在搜索之前用同義詞擴充問題。在 app.py 中:

EXPAND_QUERY = True

修改代碼會改變該執行所記錄的系統版本。再執行一次,然後把候選與基準進行比較:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

單看候選執行:citations-valid 現在 PASS,hit_rate_at_2 顯示 100.0% [75.2%, 100.0%],而閘門仍以結束代碼 3 阻止,因為 answer-floor 和 retrieval-floor 仍是 INSUFFICIENT_EVIDENCE。比較結果:

Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)

這次變更修復了它針對的案例(多答對一個,在 15 個配對案例上 +6.7 個百分點)。比較仍然無法確定候選比基準差的程度不超過邊際:15 個配對案例留下的區間寬約 67 個百分點。規劃行說明,如果差異保持不變,還需要多少配對案例才能作出決策。下一步是更大的套件,而不是換一個邊際。見比較兩次執行和比較規則。

路徑 B:帶診斷的分階段應用

分階段的檔案

路徑 B 執行 support_rag 範例,它在 RAG 評估頁面上有介紹。複製它:

oloproof init --example support_rag my-staged-rag
cd my-staged-rag
檔案它是什麼
app.pySupportRag,一個用 @rag_system 裝飾的類:retrieve(input, depth)、generate(input, context)、count_tokens(passage)
data/corpus.jsonl、data/support.jsonl知識庫,以及 13 個案例,每個都帶有 relevant 和 gold_context
oloproof.yamlsystem.rag 指向這個類,並設置 depth、top_k、token_budget、index_version
release.yaml、compare.yaml與路徑 A 相同的政策

與路徑 A 的區別在於由誰來組裝上下文。在這裡,Oloproof 呼叫 retrieve,保留前 top_k 個候選,丟棄超出 token_budget 的段落,並把其餘的傳給 generate。由於它把各階段分開,它可以分別快取它們,並用不同的上下文重新執行生成。要適配你自己的應用,替換 retrieve(呼叫你的索引,按檢索器的順序返回 Retrieval(candidates=[Passage(...)]))和 generate(用給定的段落呼叫你的模型)的函式體。把 index_version 設為一個會隨索引變化而變化的值:它是檢索身份的一部分,過時的值會讓已快取的檢索結果被複用,而當前索引已經不再返回它們了。

同樣的設定還允許使用 relevant_position 和 context_truncated 切片,以及一個覆蓋完整檢索深度的 ndcg 評估器。

執行分階段的套件

oloproof run
Gate: BLOCK (exit 3)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ PASS                  │ observed_failures_within_limit │
│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages 行是分階段系統自己的快取。結束代碼 3:沒有任何規則 FAIL,但有兩條規則缺少 PASS 所需的證據。

用標準上下文進行診斷,並與對照組並列

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…

從失敗的案例中會生成兩個子執行:一個用案例的 gold_context 代替檢索到的上下文,另一個是原樣重新執行它們的對照組。對照組讓解讀變得可靠:一個在普通重跑中就通過的案例是不穩定的,而不是被診斷出來的。無論發現什麼,diagnose 都以 0 結束;它不對發布作任何決策。

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

解讀標籤

標籤觀測到了什麼它不能確立什麼
RETRIEVAL_MISS沒有檢索到相關段落,而案例在使用標準段落時通過了檢索是唯一的問題,或者某個特定的檢索變更就能修復它
RANKED_OUT檢索到了一個相關段落,但排在 top_k 之後,而案例在使用標準段落時通過了擴大 top-k 對其他案例也會有幫助
CONTEXT_ASSEMBLY_LOSStop-k 之內的一個相關段落從上下文中被丟棄了,而案例在使用標準段落時通過了多大的預算才夠
GENERATION_FAILURE即使手握標準段落,案例仍然失敗出錯的是模型,而不是提示或期望值
UNRESOLVED無法得出任何結論:系統不是分階段的(intervention_unsupported),案例在對照組下恢復了(unstable_under_control),它沒有相關性標籤(no_relevance_labels),或證據缺失關於該案例的任何事情

每個標籤都是在對這些案例施加一種干預的情況下,失敗與某個階段之間的關聯。它不是已經證明的原因:“Implicated”和“Candidate experiment”是輸出中使用的最強的措辭,而計數只描述所選的案例(“no population claim”)。seat_count 是一個很好的提醒:它在拿到正確段落時仍然失敗,因為知識庫寫的是“five”而案例期望的是“5”,這是任何檢索變更都無法修復的。

有標準段落和沒有標準段落的案例

只有宣告瞭 expected.gold_context 的失敗案例才能被重新執行。從 seat_count 和 money_back 中移除標準段落(並移除 money_back 的相關性標籤),同一條命令會報告:

Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the context

被排除的案例會被列出,而不是被悄悄丟棄。另外請注意,移除相關性標籤對執行本身有什麼影響:hit_rate_at_2 升到了 100.0%(12 / 12 observed,1 excluded),因為檢索遺漏的那一個案例不再被量測。未標註的案例離開分母;它們不計為通過,而基於更少案例的指標可能看起來比應用的實際情況更好。先標註難的案例。

在動手之前先檢驗修復:top-k 和重排序器

另外兩種干預會用不同的設置重放記錄下來的檢索,所以不會再次呼叫檢索器:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recovered

重排序器是你編寫的一個函式 (input, candidates) -> candidates。把它保存為 app.py 旁邊的 rerank.py:

"""A candidate reranker: shorter passages first, so more of them fit the token budget."""

from oloproof import Passage


def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
    return sorted(candidates, key=lambda passage: len((passage.text or "").split()))
oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correct
Recovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recovered

兩者都沒有恢復任何案例,這正是標準上下文標籤所預測的:這裡沒有哪個失敗是因為段落剛好排在截斷線之下。每次重放都會延續標準上下文標籤,所以這些診斷可以放在一起閱讀。

執行診斷指出的實驗,並進行比較

在 oloproof.yaml 中把 token_budget 提高到 120,然後:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

每次檢索都被複用了,因為 top_k 和預算不屬於檢索的身份;只有上下文發生了變化的六個案例被重新生成。

answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

這次實驗沒有幫助:沒有一個案例改變了結論,所以診斷提出的假設在這些案例上沒有得到支援。這是一個有用的結果。下一個實驗是更小的分塊,或者那兩個上下文組裝案例的提示;seat_count 需要修復它的期望值。

故障排除

症狀原因修復
Configuration error: slice 'relevant_position' ... needs a staged system在可呼叫物件或 HTTP 系統上使用了位置切片去掉這個切片,或轉到路徑 B
執行以結束代碼 2 停止,並顯示 malformed retrieval/v1 artifact出現了模式不允許的欄位,或候選數量超過了 depth只映射文件中列出的欄位;把 depth 設為至少等於返回的數量
citations_valid 顯示 0 / 0 observed · 15 missing,它的規則是 INSUFFICIENT_EVIDENCE,原因為 no_observations適配器沒有記錄 citations/v1(或 context/v1);每個這樣的案例都是缺失,而不是通過在適配器的每條路徑上都記錄兩者,包括“沒有回答”;oloproof inspect RUN_ID --failures 會顯示每個案例的錯誤
某個檢索指標顯示很多 excluded案例沒有 expected.relevant標註它們,或者在知情的情況下接受更小的分母
diagnose 顯示 UNRESOLVED ... not staged路徑 A符合預期;要做干預請使用路徑 B
diagnose 拒絕執行並顯示 an intervention must re-execute the same system自該執行以來代碼或設定發生了變化診斷當前版本的執行,或恢復執行時的那個版本
診斷選中的案例比失敗的少失敗的案例沒有 expected.gold_context添加標準段落;被排除的案例會在輸出中列出
索引變化後檢索結果仍被複用index_version 沒有變化索引變化時修改 index_version

侷限

  • Oloproof 呼叫你的應用;它不託管、不隔離也不重置它。它的索引、快取以及它保存的任何狀態都歸你管理。
  • 在黑盒上無法進行干預:diagnose 把每個案例都標為 UNRESOLVED,並且不重新執行任何東西。
  • 干預包括標準上下文、top-k 和重排序器。沒有分塊、嵌入或提示方面的干預。
  • 診斷標籤描述的是在一種干預下、與對照組並列的所選失敗案例。它們把失敗與某個階段關聯起來;它們不證明原因,也不對未被選中的案例作任何斷言。
  • 檢索指標需要相關性標籤,診斷需要標準段落;Oloproof 不會建立這兩者。
  • 確定性範例代替了真實的檢索器和模型。在 generate 中使用真實模型或使用評判評估器,都會呼叫提供者,需要憑據,並且每個案例都要花錢。
  • 哪些功能在哪裡可用(SDK、YAML 與瀏覽器)見目前可用的功能。