指南
教程:评估一个 RAG 应用
一份针对检索增强应用的可运行演练,分两条路径:一个已有的黑盒应用,你从外部记录它的检索、上下文和引用;以及一个分阶段的应用,Oloproof 逐阶段运行它,从而让 oloproof diagnose 能在受控变更下重新执行失败的用例。两者都在本地运行,无需提供方凭据。
每一步背后的概念(阶段、相关性标签、标准上下文、四种失败标签)在 RAG 评估页面上;用例、评估器、指标、区间和门禁这些术语在核心概念中。本页是贯穿它们的动手路线。
哪条路径适合你
| 你的应用 | 路径 | 你能得到什么 | 你得不到什么 |
|---|---|---|---|
| 一次调用进,一个回答出(一个服务、一个 HTTP 端点、一条你不想拆开的框架链) | A,黑盒 | 检索指标、引用检查、依据性评判模型、门禁、比较 | 受控干预:diagnose 不重新执行任何东西 |
| 可以分别调用的检索和生成 | B,分阶段 | A 中的一切、按阶段缓存,以及带标准上下文、top-k 和重排序器并与对照组并列的 diagnose | 这三种之外的干预 |
如果不确定,就从 A 开始。它不需要修改应用,以后转到 B 时,数据集、评估器和策略都可以保留。
前提条件
- Python 3.11 或更高版本,并已安装 Oloproof(pip install oloproof)。
- 示例项目,它们随软件包一起提供:路径 A 用 blackbox_rag,路径 B 用 support_rag。把其中一个复制到新目录并在那里工作:
oloproof init --example blackbox_rag my-rag
cd my-rag下面的每条命令都在复制出的目录中运行。运行、评判结果和诊断都存储在那里的 .oloproof/ 中。
路径 A:把已有应用作为黑盒
文件
| 文件 | 它是什么 |
|---|---|
| app.py | support_api(question),代替你的应用;以及 run(case),即适配器 |
| server.py | 通过 HTTP 提供的同一个应用,用于下面的 HTTP 变体 |
| data/corpus.jsonl | 应用所搜索的、包含 14 个段落的知识库 |
| data/support.jsonl | 15 个用例:13 个带有相关性标签和标准段落,2 个两者都没有 |
| oloproof.yaml | 测试套件:数据集、系统、评估器、切片 |
| oloproof.http.yaml | 针对 HTTP 服务器的同一个测试套件 |
| release.yaml | 单次运行的发布策略 |
| compare.yaml | 把候选运行与基线进行比较的策略 |
应用返回什么
support_api 的行为就像一个你已经拥有的应用:它进行搜索,从符合字数预算的最佳来源构建提示,作答并引用。它的响应已经带有它所做的事:
{
"answer": "Team plans include five seats.",
"cited": ["kb-03"],
"sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
"prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
"skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}你的应用的字段名会有所不同。重要的是,它能针对每个问题告诉你它检索到的排序后的来源、到达模型的来源,以及它引用的来源。如果做不到,先把这些加入它的响应或日志:Oloproof 度量的是被记录下来的内容,绝不会从回答中推断检索。
适配器
run 原样调用应用,并把响应映射为三种有类型的产物,也就是检索和引用评估器所读取的记录:
@system(
name="support-rag-blackbox",
version="tutorial",
records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
response = support_api(str(case["question"]))
recorder = current_case()
recorder.retrieval(
Retrieval(
query=case["question"],
depth=SEARCH_DEPTH,
candidates=tuple(
Passage(doc_id=s["id"], score=s["score"], text=s["text"])
for s in response["sources"]
),
)
)
recorder.context(
Context(
items=tuple(
ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
for i in response["prompt_sources"]
),
dropped=tuple(
DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
for i in response["skipped"]
),
token_budget=PROMPT_WORD_BUDGET,
)
)
recorder.citations(response["cited"])
return {"answer": response["answer"], "citations": response["cited"]}| 产物 | 结构 | 读取者 |
|---|---|---|
| retrieval/v1 | query、depth,以及按检索器返回顺序排列的 candidates,每个都是一个 Passage(doc_id, chunk_id, score, text) | hit_rate、recall、mrr、ndcg |
| context/v1 | 到达模型的 items(doc_id、position、tokens、text),带有 top_k 或 token_budget 这一 reason 的 dropped 项,以及 token_budget | citation_validity、groundedness_judge、citation_support_judge |
| citations/v1 | ids,每个是一个 doc_id 或 doc_id#chunk_id | citation_validity、citation_support_judge |
Oloproof 按给定的位置记录,绝不重新排序。格式错误的产物会让运行以退出码 2 停止,而不会被存储。case 是用例的 input 对象,所以 case["question"] 就是数据集中的问题。
要使用你自己的应用,把 support_api 的函数体替换为对它的调用(一次 SDK 调用、一个 HTTP 请求),并保留 run。在 oloproof.yaml 中把 system.callable 以 module:function 的形式指向它。
HTTP 变体
HTTP 系统无法调用记录器,所以改由它的响应携带证据,并且已经是上面的三种结构,配置则指明证据所在的位置:
system:
name: support-rag-http
version: tutorial
http:
url: http://127.0.0.1:8766/answer
output_path: result
artifacts:
retrieval/v1: evidence.retrieval
context/v1: evidence.context
citations/v1: evidence.citationsserver.py 提供的正是这些。启动它,然后针对它运行:
python server.py 8766
oloproof run --config oloproof.http.yaml用例输入作为 JSON 请求体发送。output_path 从响应中取出输出,每个 artifacts 条目把一个点分路径记录为该类型;缺失或格式错误的字段会让运行以退出码 2 停止。结果与下面的可调用对象路径完全相同。在你自己的服务中,证据对象通常是一个你为评估流量启用的调试字段。
用例声明了什么
{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}- expected.relevant 列出回答该问题的段落。检索指标读取它。没有它的用例,比如 office_hours,会以 no_relevance_labels 被排除在这些指标之外:它离开分母,而不是被计为通过或失败。
- expected.gold_context 是段落文本本身。路径 A 从不使用它;路径 B 在诊断时用它替换检索到的上下文。
在实践中,未标注的用例很常见,因为标注相关性需要工作量。它们仍然计入回答和引用检查。
选择评估器
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: hit_rate, k: 2}
- {type: recall, k: 2}
- {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4- contains 检查回答是否包含预期文本。它是任务检查:用户是否得到了正确的答案。措辞多变时,改用精确匹配或评分细则评判模型。
- k: 2 下的 hit_rate 和 recall 在应用实际放进提示的深度上度量检索。在模型从未看到的深度上的检索指标,描述的是索引,而不是应用。
- citation_validity 检查每个被引用的 id 是否指向一个到达了模型的段落;require_citations: true 还会让什么都不引用的回答失败。
- groundedness_judge 和 citation_support_judge(可选)询问模型回答是否得到上下文的支持。它们需要一个提供方、一个模型和环境变量中的凭据,并且每个用例都要花钱;评判模型在获准做门禁之前必须达到什么要求,见评判器。
这里无法使用 relevant_position 和 context_truncated 切片:它们把位置与应用的 top-k 进行比较,而只有分阶段的系统才会声明 top-k。请求它们会让运行停止,并显示 slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system。
发布策略
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answer-floor
metric: answer_correct
min: 0.70
- id: retrieval-floor
metric: hit_rate_at_2
min: 0.80
- id: citations-valid
metric: citations_valid
kind: observed_count
max_failures: 0min 规则只有在整个区间都越过下限时才通过,在整个区间都低于下限时失败,否则为 INSUFFICIENT_EVIDENCE。observed_count 规则依据实际运行的用例作出决策,没有区间:“这个测试套件中没有无效引用”。见门禁。
运行
oloproof runRun run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 73.3% │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3% │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 miss如何阅读:
- Gate: BLOCK (exit 1):有一条规则 FAIL 了。退出码 1 表示出现 FAIL;退出码 3 表示门禁在没有 FAIL 的情况下阻止了发布(这里会是 INSUFFICIENT_EVIDENCE);退出码 0 表示没有出现任何策略要阻止的情况。[DECIDED/COMPLETE] 是执行状态:每个用例都运行了。
- citations-valid FAIL:有一个回答什么都没有引用,而 require_citations 把这算作无效。
- answer-floor 是 INSUFFICIENT_EVIDENCE 而不是 PASS,尽管 73.3% 高于 70%:只有 15 个用例时,区间向下延伸到 44.8%,所以证据无法证明达到了下限。
- hit_rate_at_2 显示 2 excluded:就是那两个未标注的用例。它的分母是 13,而不是 15。
- 后面的 Slices 表是探索性的,从不用于门禁;低于 min_slice_support 的切片不显示区间。
查看失败项
运行 id 在运行输出的第一行。
oloproof inspect RUN_ID --failures4 of 15 cases failed, errored or did not finish
refund_review
output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
answer_correct: failed
money_back
output: {"answer": "I could not find that in the knowledge base.", "citations": []}
answer_correct: failed
hit_rate_at_2: failed
recall_at_2: failed
citations_valid: failed
security_review
output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
answer_correct: failed
seat_count
output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
answer_correct: failedoloproof inspect RUN_ID --case refund_review 打印一个用例的输入、预期值、输出和每个评判结果。记录下来的产物在导出的包中:
oloproof export RUN_ID.oloproof/bundles/RUN_ID/cases.jsonl 的每一行是一个用例的记录;它的 artifacts 字段保存了被记录的内容。对于 money_back,该字段内容为:
{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}仅凭记录下来的证据来解读这四个失败:
| 用例 | 记录显示了什么 | 有意义的下一步 |
|---|---|---|
| money_back | 检索什么都没有返回:问题与退款段落没有任何共同的词 | 查询改写或同义词,用 hit_rate_at_2 来度量 |
| refund_review、security_review | hit_rate_at_2 通过了,但回答来自另一个段落 | 查看 context/v1:相关段落是否因为预算而被丢弃了? |
| seat_count | 正确的段落被检索到、被保留并被引用;回答说“five”,用例期望的是“5” | 修复期望值或回答格式,而不是检索 |
这张表是你对记录的解读。它是失败与某个阶段之间的关联,而不是已经证明的原因:没有任何东西在改变该阶段后重新执行过这个用例。
diagnose 对黑盒做了什么
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc每个用例都是 UNRESOLVED,原因为 intervention_unsupported。Oloproof 无法把标准段落交给黑盒来代替它自己的检索,所以它不会假装能做到。受控干预需要路径 B。
做一次候选变更并比较
记录表明 money_back 在检索阶段失败。候选变更会在搜索之前用同义词扩展问题。在 app.py 中:
EXPAND_QUERY = True修改代码会改变该运行所记录的系统版本。再运行一次,然后把候选与基线进行比较:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml单看候选运行:citations-valid 现在 PASS,hit_rate_at_2 显示 100.0% [75.2%, 100.0%],而门禁仍以退出码 3 阻止,因为 answer-floor 和 retrieval-floor 仍是 INSUFFICIENT_EVIDENCE。比较结果:
Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)这次变更修复了它针对的用例(多答对一个,在 15 个配对用例上 +6.7 个百分点)。比较仍然无法确定候选比基线差的程度不超过边际:15 个配对用例留下的区间宽约 67 个百分点。规划行说明,如果差异保持不变,还需要多少配对用例才能作出决策。下一步是更大的测试套件,而不是换一个边际。见比较两次运行和比较规则。
路径 B:带诊断的分阶段应用
分阶段的文件
路径 B 运行 support_rag 示例,它在 RAG 评估页面上有介绍。复制它:
oloproof init --example support_rag my-staged-rag
cd my-staged-rag| 文件 | 它是什么 |
|---|---|
| app.py | SupportRag,一个用 @rag_system 装饰的类:retrieve(input, depth)、generate(input, context)、count_tokens(passage) |
| data/corpus.jsonl、data/support.jsonl | 知识库,以及 13 个用例,每个都带有 relevant 和 gold_context |
| oloproof.yaml | system.rag 指向这个类,并设置 depth、top_k、token_budget、index_version |
| release.yaml、compare.yaml | 与路径 A 相同的策略 |
与路径 A 的区别在于由谁来组装上下文。在这里,Oloproof 调用 retrieve,保留前 top_k 个候选,丢弃超出 token_budget 的段落,并把其余的传给 generate。由于它把各阶段分开,它可以分别缓存它们,并用不同的上下文重新执行生成。要适配你自己的应用,替换 retrieve(调用你的索引,按检索器的顺序返回 Retrieval(candidates=[Passage(...)]))和 generate(用给定的段落调用你的模型)的函数体。把 index_version 设为一个会随索引变化而变化的值:它是检索身份的一部分,过时的值会让已缓存的检索结果被复用,而当前索引已经不再返回它们了。
同样的配置还允许使用 relevant_position 和 context_truncated 切片,以及一个覆盖完整检索深度的 ndcg 评估器。
运行分阶段的测试套件
oloproof runGate: BLOCK (exit 3)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ PASS │ observed_failures_within_limit │
│ answer_correct │ 69.2% │ [38.5%, 91.0%] │ 9 / 13 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ ndcg_at_6 │ 0.866 │ [0.506, 0.990] │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0% │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 missStages 行是分阶段系统自己的缓存。退出码 3:没有任何规则 FAIL,但有两条规则缺少 PASS 所需的证据。
用标准上下文进行诊断,并与对照组并列
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…从失败的用例中会生成两个子运行:一个用用例的 gold_context 代替检索到的上下文,另一个是原样重新执行它们的对照组。对照组让解读变得可靠:一个在普通重跑中就通过的用例是不稳定的,而不是被诊断出来的。无论发现什么,diagnose 都以 0 退出;它不对发布作任何决策。
oloproof inspect DIAGNOSIS_IDmoney_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2解读标签
| 标签 | 观测到了什么 | 它不能确立什么 |
|---|---|---|
| RETRIEVAL_MISS | 没有检索到相关段落,而用例在使用标准段落时通过了 | 检索是唯一的问题,或者某个特定的检索变更就能修复它 |
| RANKED_OUT | 检索到了一个相关段落,但排在 top_k 之后,而用例在使用标准段落时通过了 | 扩大 top-k 对其他用例也会有帮助 |
| CONTEXT_ASSEMBLY_LOSS | top-k 之内的一个相关段落从上下文中被丢弃了,而用例在使用标准段落时通过了 | 多大的预算才够 |
| GENERATION_FAILURE | 即使手握标准段落,用例仍然失败 | 出错的是模型,而不是提示或期望值 |
| UNRESOLVED | 无法得出任何结论:系统不是分阶段的(intervention_unsupported),用例在对照组下恢复了(unstable_under_control),它没有相关性标签(no_relevance_labels),或证据缺失 | 关于该用例的任何事情 |
每个标签都是在对这些用例施加一种干预的情况下,失败与某个阶段之间的关联。它不是已经证明的原因:“Implicated”和“Candidate experiment”是输出中使用的最强的措辞,而计数只描述所选的用例(“no population claim”)。seat_count 是一个很好的提醒:它在拿到正确段落时仍然失败,因为知识库写的是“five”而用例期望的是“5”,这是任何检索变更都无法修复的。
有标准段落和没有标准段落的用例
只有声明了 expected.gold_context 的失败用例才能被重新执行。从 seat_count 和 money_back 中移除标准段落(并移除 money_back 的相关性标签),同一条命令会报告:
Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the context被排除的用例会被列出,而不是被悄悄丢弃。另外请注意,移除相关性标签对运行本身有什么影响:hit_rate_at_2 升到了 100.0%(12 / 12 observed,1 excluded),因为检索遗漏的那一个用例不再被度量。未标注的用例离开分母;它们不计为通过,而基于更少用例的指标可能看起来比应用的实际情况更好。先标注难的用例。
在动手之前先检验修复:top-k 和重排序器
另外两种干预会用不同的设置重放记录下来的检索,所以不会再次调用检索器:
oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correctRecovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recovered重排序器是你编写的一个函数 (input, candidates) -> candidates。把它保存为 app.py 旁边的 rerank.py:
"""A candidate reranker: shorter passages first, so more of them fit the token budget."""
from oloproof import Passage
def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
return sorted(candidates, key=lambda passage: len((passage.text or "").split()))oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correctRecovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recovered两者都没有恢复任何用例,这正是标准上下文标签所预测的:这里没有哪个失败是因为段落刚好排在截断线之下。每次重放都会延续标准上下文标签,所以这些诊断可以放在一起阅读。
运行诊断指出的实验,并进行比较
在 oloproof.yaml 中把 token_budget 提高到 120,然后:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlStages: retrieve 13 hit/0 miss; generate 7 hit/6 miss每次检索都被复用了,因为 top_k 和预算不属于检索的身份;只有上下文发生了变化的六个用例被重新生成。
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
Gate: BLOCK (exit 3)这次实验没有帮助:没有一个用例改变了结论,所以诊断提出的假设在这些用例上没有得到支持。这是一个有用的结果。下一个实验是更小的分块,或者那两个上下文组装用例的提示;seat_count 需要修复它的期望值。
故障排除
| 症状 | 原因 | 修复 |
|---|---|---|
| Configuration error: slice 'relevant_position' ... needs a staged system | 在可调用对象或 HTTP 系统上使用了位置切片 | 去掉这个切片,或转到路径 B |
| 运行以退出码 2 停止,并显示 malformed retrieval/v1 artifact | 出现了模式不允许的字段,或候选数量超过了 depth | 只映射文档中列出的字段;把 depth 设为至少等于返回的数量 |
| citations_valid 显示 0 / 0 observed · 15 missing,它的规则是 INSUFFICIENT_EVIDENCE,原因为 no_observations | 适配器没有记录 citations/v1(或 context/v1);每个这样的用例都是缺失,而不是通过 | 在适配器的每条路径上都记录两者,包括“没有回答”;oloproof inspect RUN_ID --failures 会显示每个用例的错误 |
| 某个检索指标显示很多 excluded | 用例没有 expected.relevant | 标注它们,或者在知情的情况下接受更小的分母 |
| diagnose 显示 UNRESOLVED ... not staged | 路径 A | 符合预期;要做干预请使用路径 B |
| diagnose 拒绝执行并显示 an intervention must re-execute the same system | 自该运行以来代码或配置发生了变化 | 诊断当前版本的运行,或恢复运行时的那个版本 |
| 诊断选中的用例比失败的少 | 失败的用例没有 expected.gold_context | 添加标准段落;被排除的用例会在输出中列出 |
| 索引变化后检索结果仍被复用 | index_version 没有变化 | 索引变化时修改 index_version |
局限
- Oloproof 调用你的应用;它不托管、不隔离也不重置它。它的索引、缓存以及它保存的任何状态都归你管理。
- 在黑盒上无法进行干预:diagnose 把每个用例都标为 UNRESOLVED,并且不重新执行任何东西。
- 干预包括标准上下文、top-k 和重排序器。没有分块、嵌入或提示方面的干预。
- 诊断标签描述的是在一种干预下、与对照组并列的所选失败用例。它们把失败与某个阶段关联起来;它们不证明原因,也不对未被选中的用例作任何断言。
- 检索指标需要相关性标签,诊断需要标准段落;Oloproof 不会创建这两者。
- 确定性示例代替了真实的检索器和模型。在 generate 中使用真实模型或使用评判评估器,都会调用提供方,需要凭据,并且每个用例都要花钱。
- 哪些功能在哪里可用(SDK、YAML 与浏览器)见目前可用的功能。