跳到主要内容

指南

教程:评估一个 RAG 应用

一份针对检索增强应用的可运行演练,分两条路径:一个已有的黑盒应用,你从外部记录它的检索、上下文和引用;以及一个分阶段的应用,Oloproof 逐阶段运行它,从而让 oloproof diagnose 能在受控变更下重新执行失败的用例。两者都在本地运行,无需提供方凭据。

每一步背后的概念(阶段、相关性标签、标准上下文、四种失败标签)在 RAG 评估页面上;用例、评估器、指标、区间和门禁这些术语在核心概念中。本页是贯穿它们的动手路线。

哪条路径适合你

你的应用路径你能得到什么你得不到什么
一次调用进,一个回答出(一个服务、一个 HTTP 端点、一条你不想拆开的框架链)A,黑盒检索指标、引用检查、依据性评判模型、门禁、比较受控干预:diagnose 不重新执行任何东西
可以分别调用的检索和生成B,分阶段A 中的一切、按阶段缓存,以及带标准上下文、top-k 和重排序器并与对照组并列的 diagnose这三种之外的干预

如果不确定,就从 A 开始。它不需要修改应用,以后转到 B 时,数据集、评估器和策略都可以保留。

前提条件

  • Python 3.11 或更高版本,并已安装 Oloproof(pip install oloproof)。
  • 示例项目,它们随软件包一起提供:路径 A 用 blackbox_rag,路径 B 用 support_rag。把其中一个复制到新目录并在那里工作:
oloproof init --example blackbox_rag my-rag
cd my-rag

下面的每条命令都在复制出的目录中运行。运行、评判结果和诊断都存储在那里的 .oloproof/ 中。

路径 A:把已有应用作为黑盒

文件

文件它是什么
app.pysupport_api(question),代替你的应用;以及 run(case),即适配器
server.py通过 HTTP 提供的同一个应用,用于下面的 HTTP 变体
data/corpus.jsonl应用所搜索的、包含 14 个段落的知识库
data/support.jsonl15 个用例:13 个带有相关性标签和标准段落,2 个两者都没有
oloproof.yaml测试套件:数据集、系统、评估器、切片
oloproof.http.yaml针对 HTTP 服务器的同一个测试套件
release.yaml单次运行的发布策略
compare.yaml把候选运行与基线进行比较的策略

应用返回什么

support_api 的行为就像一个你已经拥有的应用:它进行搜索,从符合字数预算的最佳来源构建提示,作答并引用。它的响应已经带有它所做的事:

{
  "answer": "Team plans include five seats.",
  "cited": ["kb-03"],
  "sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
  "prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
  "skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}

你的应用的字段名会有所不同。重要的是,它能针对每个问题告诉你它检索到的排序后的来源、到达模型的来源,以及它引用的来源。如果做不到,先把这些加入它的响应或日志:Oloproof 度量的是被记录下来的内容,绝不会从回答中推断检索。

适配器

run 原样调用应用,并把响应映射为三种有类型的产物,也就是检索和引用评估器所读取的记录:

@system(
    name="support-rag-blackbox",
    version="tutorial",
    records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
    response = support_api(str(case["question"]))
    recorder = current_case()
    recorder.retrieval(
        Retrieval(
            query=case["question"],
            depth=SEARCH_DEPTH,
            candidates=tuple(
                Passage(doc_id=s["id"], score=s["score"], text=s["text"])
                for s in response["sources"]
            ),
        )
    )
    recorder.context(
        Context(
            items=tuple(
                ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
                for i in response["prompt_sources"]
            ),
            dropped=tuple(
                DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
                for i in response["skipped"]
            ),
            token_budget=PROMPT_WORD_BUDGET,
        )
    )
    recorder.citations(response["cited"])
    return {"answer": response["answer"], "citations": response["cited"]}
产物结构读取者
retrieval/v1query、depth,以及按检索器返回顺序排列的 candidates,每个都是一个 Passage(doc_id, chunk_id, score, text)hit_rate、recall、mrr、ndcg
context/v1到达模型的 items(doc_id、position、tokens、text),带有 top_k 或 token_budget 这一 reason 的 dropped 项,以及 token_budgetcitation_validity、groundedness_judge、citation_support_judge
citations/v1ids,每个是一个 doc_id 或 doc_id#chunk_idcitation_validity、citation_support_judge

Oloproof 按给定的位置记录,绝不重新排序。格式错误的产物会让运行以退出码 2 停止,而不会被存储。case 是用例的 input 对象,所以 case["question"] 就是数据集中的问题。

要使用你自己的应用,把 support_api 的函数体替换为对它的调用(一次 SDK 调用、一个 HTTP 请求),并保留 run。在 oloproof.yaml 中把 system.callable 以 module:function 的形式指向它。

HTTP 变体

HTTP 系统无法调用记录器,所以改由它的响应携带证据,并且已经是上面的三种结构,配置则指明证据所在的位置:

system:
  name: support-rag-http
  version: tutorial
  http:
    url: http://127.0.0.1:8766/answer
    output_path: result
    artifacts:
      retrieval/v1: evidence.retrieval
      context/v1: evidence.context
      citations/v1: evidence.citations

server.py 提供的正是这些。启动它,然后针对它运行:

python server.py 8766
oloproof run --config oloproof.http.yaml

用例输入作为 JSON 请求体发送。output_path 从响应中取出输出,每个 artifacts 条目把一个点分路径记录为该类型;缺失或格式错误的字段会让运行以退出码 2 停止。结果与下面的可调用对象路径完全相同。在你自己的服务中,证据对象通常是一个你为评估流量启用的调试字段。

用例声明了什么

{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}
  • expected.relevant 列出回答该问题的段落。检索指标读取它。没有它的用例,比如 office_hours,会以 no_relevance_labels 被排除在这些指标之外:它离开分母,而不是被计为通过或失败。
  • expected.gold_context 是段落文本本身。路径 A 从不使用它;路径 B 在诊断时用它替换检索到的上下文。

在实践中,未标注的用例很常见,因为标注相关性需要工作量。它们仍然计入回答和引用检查。

选择评估器

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4
  • contains 检查回答是否包含预期文本。它是任务检查:用户是否得到了正确的答案。措辞多变时,改用精确匹配或评分细则评判模型。
  • k: 2 下的 hit_rate 和 recall 在应用实际放进提示的深度上度量检索。在模型从未看到的深度上的检索指标,描述的是索引,而不是应用。
  • citation_validity 检查每个被引用的 id 是否指向一个到达了模型的段落;require_citations: true 还会让什么都不引用的回答失败。
  • groundedness_judge 和 citation_support_judge(可选)询问模型回答是否得到上下文的支持。它们需要一个提供方、一个模型和环境变量中的凭据,并且每个用例都要花钱;评判模型在获准做门禁之前必须达到什么要求,见评判器。

这里无法使用 relevant_position 和 context_truncated 切片:它们把位置与应用的 top-k 进行比较,而只有分阶段的系统才会声明 top-k。请求它们会让运行停止,并显示 slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system。

发布策略

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: retrieval-floor
    metric: hit_rate_at_2
    min: 0.80
  - id: citations-valid
    metric: citations_valid
    kind: observed_count
    max_failures: 0

min 规则只有在整个区间都越过下限时才通过,在整个区间都低于下限时失败,否则为 INSUFFICIENT_EVIDENCE。observed_count 规则依据实际运行的用例作出决策,没有区间:“这个测试套件中没有无效引用”。见门禁。

运行

oloproof run
Run run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct  │ 73.3%    │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3%    │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 miss

如何阅读:

  • Gate: BLOCK (exit 1):有一条规则 FAIL 了。退出码 1 表示出现 FAIL;退出码 3 表示门禁在没有 FAIL 的情况下阻止了发布(这里会是 INSUFFICIENT_EVIDENCE);退出码 0 表示没有出现任何策略要阻止的情况。[DECIDED/COMPLETE] 是执行状态:每个用例都运行了。
  • citations-valid FAIL:有一个回答什么都没有引用,而 require_citations 把这算作无效。
  • answer-floor 是 INSUFFICIENT_EVIDENCE 而不是 PASS,尽管 73.3% 高于 70%:只有 15 个用例时,区间向下延伸到 44.8%,所以证据无法证明达到了下限。
  • hit_rate_at_2 显示 2 excluded:就是那两个未标注的用例。它的分母是 13,而不是 15。
  • 后面的 Slices 表是探索性的,从不用于门禁;低于 min_slice_support 的切片不显示区间。

查看失败项

运行 id 在运行输出的第一行。

oloproof inspect RUN_ID --failures
4 of 15 cases failed, errored or did not finish

refund_review
  output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
  answer_correct: failed

money_back
  output: {"answer": "I could not find that in the knowledge base.", "citations": []}
  answer_correct: failed
  hit_rate_at_2: failed
  recall_at_2: failed
  citations_valid: failed

security_review
  output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
  answer_correct: failed

seat_count
  output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
  answer_correct: failed

oloproof inspect RUN_ID --case refund_review 打印一个用例的输入、预期值、输出和每个评判结果。记录下来的产物在导出的包中:

oloproof export RUN_ID

.oloproof/bundles/RUN_ID/cases.jsonl 的每一行是一个用例的记录;它的 artifacts 字段保存了被记录的内容。对于 money_back,该字段内容为:

{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}

仅凭记录下来的证据来解读这四个失败:

用例记录显示了什么有意义的下一步
money_back检索什么都没有返回:问题与退款段落没有任何共同的词查询改写或同义词,用 hit_rate_at_2 来度量
refund_review、security_reviewhit_rate_at_2 通过了,但回答来自另一个段落查看 context/v1:相关段落是否因为预算而被丢弃了?
seat_count正确的段落被检索到、被保留并被引用;回答说“five”,用例期望的是“5”修复期望值或回答格式,而不是检索

这张表是你对记录的解读。它是失败与某个阶段之间的关联,而不是已经证明的原因:没有任何东西在改变该阶段后重新执行过这个用例。

diagnose 对黑盒做了什么

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc

每个用例都是 UNRESOLVED,原因为 intervention_unsupported。Oloproof 无法把标准段落交给黑盒来代替它自己的检索,所以它不会假装能做到。受控干预需要路径 B。

做一次候选变更并比较

记录表明 money_back 在检索阶段失败。候选变更会在搜索之前用同义词扩展问题。在 app.py 中:

EXPAND_QUERY = True

修改代码会改变该运行所记录的系统版本。再运行一次,然后把候选与基线进行比较:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

单看候选运行:citations-valid 现在 PASS,hit_rate_at_2 显示 100.0% [75.2%, 100.0%],而门禁仍以退出码 3 阻止,因为 answer-floor 和 retrieval-floor 仍是 INSUFFICIENT_EVIDENCE。比较结果:

Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)

这次变更修复了它针对的用例(多答对一个,在 15 个配对用例上 +6.7 个百分点)。比较仍然无法确定候选比基线差的程度不超过边际:15 个配对用例留下的区间宽约 67 个百分点。规划行说明,如果差异保持不变,还需要多少配对用例才能作出决策。下一步是更大的测试套件,而不是换一个边际。见比较两次运行和比较规则。

路径 B:带诊断的分阶段应用

分阶段的文件

路径 B 运行 support_rag 示例,它在 RAG 评估页面上有介绍。复制它:

oloproof init --example support_rag my-staged-rag
cd my-staged-rag
文件它是什么
app.pySupportRag,一个用 @rag_system 装饰的类:retrieve(input, depth)、generate(input, context)、count_tokens(passage)
data/corpus.jsonl、data/support.jsonl知识库,以及 13 个用例,每个都带有 relevant 和 gold_context
oloproof.yamlsystem.rag 指向这个类,并设置 depth、top_k、token_budget、index_version
release.yaml、compare.yaml与路径 A 相同的策略

与路径 A 的区别在于由谁来组装上下文。在这里,Oloproof 调用 retrieve,保留前 top_k 个候选,丢弃超出 token_budget 的段落,并把其余的传给 generate。由于它把各阶段分开,它可以分别缓存它们,并用不同的上下文重新执行生成。要适配你自己的应用,替换 retrieve(调用你的索引,按检索器的顺序返回 Retrieval(candidates=[Passage(...)]))和 generate(用给定的段落调用你的模型)的函数体。把 index_version 设为一个会随索引变化而变化的值:它是检索身份的一部分,过时的值会让已缓存的检索结果被复用,而当前索引已经不再返回它们了。

同样的配置还允许使用 relevant_position 和 context_truncated 切片,以及一个覆盖完整检索深度的 ndcg 评估器。

运行分阶段的测试套件

oloproof run
Gate: BLOCK (exit 3)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ PASS                  │ observed_failures_within_limit │
│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

Stages 行是分阶段系统自己的缓存。退出码 3:没有任何规则 FAIL,但有两条规则缺少 PASS 所需的证据。

用标准上下文进行诊断,并与对照组并列

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…

从失败的用例中会生成两个子运行:一个用用例的 gold_context 代替检索到的上下文,另一个是原样重新执行它们的对照组。对照组让解读变得可靠:一个在普通重跑中就通过的用例是不稳定的,而不是被诊断出来的。无论发现什么,diagnose 都以 0 退出;它不对发布作任何决策。

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

解读标签

标签观测到了什么它不能确立什么
RETRIEVAL_MISS没有检索到相关段落,而用例在使用标准段落时通过了检索是唯一的问题,或者某个特定的检索变更就能修复它
RANKED_OUT检索到了一个相关段落,但排在 top_k 之后,而用例在使用标准段落时通过了扩大 top-k 对其他用例也会有帮助
CONTEXT_ASSEMBLY_LOSStop-k 之内的一个相关段落从上下文中被丢弃了,而用例在使用标准段落时通过了多大的预算才够
GENERATION_FAILURE即使手握标准段落,用例仍然失败出错的是模型,而不是提示或期望值
UNRESOLVED无法得出任何结论:系统不是分阶段的(intervention_unsupported),用例在对照组下恢复了(unstable_under_control),它没有相关性标签(no_relevance_labels),或证据缺失关于该用例的任何事情

每个标签都是在对这些用例施加一种干预的情况下,失败与某个阶段之间的关联。它不是已经证明的原因:“Implicated”和“Candidate experiment”是输出中使用的最强的措辞,而计数只描述所选的用例(“no population claim”)。seat_count 是一个很好的提醒:它在拿到正确段落时仍然失败,因为知识库写的是“five”而用例期望的是“5”,这是任何检索变更都无法修复的。

有标准段落和没有标准段落的用例

只有声明了 expected.gold_context 的失败用例才能被重新执行。从 seat_count 和 money_back 中移除标准段落(并移除 money_back 的相关性标签),同一条命令会报告:

Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the context

被排除的用例会被列出,而不是被悄悄丢弃。另外请注意,移除相关性标签对运行本身有什么影响:hit_rate_at_2 升到了 100.0%(12 / 12 observed,1 excluded),因为检索遗漏的那一个用例不再被度量。未标注的用例离开分母;它们不计为通过,而基于更少用例的指标可能看起来比应用的实际情况更好。先标注难的用例。

在动手之前先检验修复:top-k 和重排序器

另外两种干预会用不同的设置重放记录下来的检索,所以不会再次调用检索器:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recovered

重排序器是你编写的一个函数 (input, candidates) -> candidates。把它保存为 app.py 旁边的 rerank.py:

"""A candidate reranker: shorter passages first, so more of them fit the token budget."""

from oloproof import Passage


def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
    return sorted(candidates, key=lambda passage: len((passage.text or "").split()))
oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correct
Recovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recovered

两者都没有恢复任何用例,这正是标准上下文标签所预测的:这里没有哪个失败是因为段落刚好排在截断线之下。每次重放都会延续标准上下文标签,所以这些诊断可以放在一起阅读。

运行诊断指出的实验,并进行比较

在 oloproof.yaml 中把 token_budget 提高到 120,然后:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

每次检索都被复用了,因为 top_k 和预算不属于检索的身份;只有上下文发生了变化的六个用例被重新生成。

answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

这次实验没有帮助:没有一个用例改变了结论,所以诊断提出的假设在这些用例上没有得到支持。这是一个有用的结果。下一个实验是更小的分块,或者那两个上下文组装用例的提示;seat_count 需要修复它的期望值。

故障排除

症状原因修复
Configuration error: slice 'relevant_position' ... needs a staged system在可调用对象或 HTTP 系统上使用了位置切片去掉这个切片,或转到路径 B
运行以退出码 2 停止,并显示 malformed retrieval/v1 artifact出现了模式不允许的字段,或候选数量超过了 depth只映射文档中列出的字段;把 depth 设为至少等于返回的数量
citations_valid 显示 0 / 0 observed · 15 missing,它的规则是 INSUFFICIENT_EVIDENCE,原因为 no_observations适配器没有记录 citations/v1(或 context/v1);每个这样的用例都是缺失,而不是通过在适配器的每条路径上都记录两者,包括“没有回答”;oloproof inspect RUN_ID --failures 会显示每个用例的错误
某个检索指标显示很多 excluded用例没有 expected.relevant标注它们,或者在知情的情况下接受更小的分母
diagnose 显示 UNRESOLVED ... not staged路径 A符合预期;要做干预请使用路径 B
diagnose 拒绝执行并显示 an intervention must re-execute the same system自该运行以来代码或配置发生了变化诊断当前版本的运行,或恢复运行时的那个版本
诊断选中的用例比失败的少失败的用例没有 expected.gold_context添加标准段落;被排除的用例会在输出中列出
索引变化后检索结果仍被复用index_version 没有变化索引变化时修改 index_version

局限

  • Oloproof 调用你的应用;它不托管、不隔离也不重置它。它的索引、缓存以及它保存的任何状态都归你管理。
  • 在黑盒上无法进行干预:diagnose 把每个用例都标为 UNRESOLVED,并且不重新执行任何东西。
  • 干预包括标准上下文、top-k 和重排序器。没有分块、嵌入或提示方面的干预。
  • 诊断标签描述的是在一种干预下、与对照组并列的所选失败用例。它们把失败与某个阶段关联起来;它们不证明原因,也不对未被选中的用例作任何断言。
  • 检索指标需要相关性标签,诊断需要标准段落;Oloproof 不会创建这两者。
  • 确定性示例代替了真实的检索器和模型。在 generate 中使用真实模型或使用评判评估器,都会调用提供方,需要凭据,并且每个用例都要花钱。
  • 哪些功能在哪里可用(SDK、YAML 与浏览器)见目前可用的功能。