Skip to content

Guides

Tutorial: evaluate a RAG application

A runnable walkthrough for a retrieval-augmented application, in two paths: an existing black-box application whose retrieval, context and citations you record from the outside, and a staged application Oloproof runs stage by stage so oloproof diagnose can re-execute failed cases under controlled changes. Both run locally with no provider credentials.

The concepts behind each step (stages, relevance labels, gold context, the four failure labels) are on the RAG evaluation page; the terms case, evaluator, metric, interval and gate are in Core concepts. This page is the hands-on route through them.

Which path is yours

Your applicationPathWhat you getWhat you do not get
One call in, one answer out (a service, an HTTP endpoint, a framework chain you do not want to split)A, black boxRetrieval metrics, citation checks, grounding judges, gating, comparisonControlled interventions: diagnose re-executes nothing
Retrieval and generation you can call separatelyB, stagedEverything in A, per-stage caching, and diagnose with gold context, top-k and a reranker beside a controlInterventions other than those three

Start with A if you are unsure. It needs no change to the application, and moving to B later keeps the dataset, evaluators and policy.

Prerequisites

  • Python 3.11 or later, and Oloproof installed (pip install oloproof).
  • The example projects, which ship with the package: blackbox_rag for path A and support_rag for path B. Copy one into a new directory and work there:
oloproof init --example blackbox_rag my-rag
cd my-rag

Every command below runs from inside the copied directory. Runs, judgments and diagnoses are stored in .oloproof/ there.

Path A: an existing application as a black box

The files

FileWhat it is
app.pysupport_api(question), standing in for your application, and run(case), the adapter
server.pyThe same application over HTTP, for the HTTP variant below
data/corpus.jsonlThe 14-passage knowledge base the application searches
data/support.jsonl15 cases: 13 with relevance labels and gold passages, 2 without either
oloproof.yamlThe suite: dataset, system, evaluators, slices
oloproof.http.yamlThe same suite against the HTTP server
release.yamlThe release policy for a single run
compare.yamlThe policy for comparing a candidate run with a baseline

What the application returns

support_api behaves like an application you already have: it searches, builds a prompt from the best sources that fit a word budget, answers and cites. Its response already carries what it did:

{
  "answer": "Team plans include five seats.",
  "cited": ["kb-03"],
  "sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
  "prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
  "skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}

Your application's field names will differ. What matters is that it can tell you, per question, the ranked sources it retrieved, the ones that reached the model, and the ones it cited. If it cannot, add those to its response or its logs first: Oloproof measures what is recorded and never infers retrieval from an answer.

The adapter

run calls the application unchanged and maps the response onto three typed artifacts, the records the retrieval and citation evaluators read:

@system(
    name="support-rag-blackbox",
    version="tutorial",
    records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
    response = support_api(str(case["question"]))
    recorder = current_case()
    recorder.retrieval(
        Retrieval(
            query=case["question"],
            depth=SEARCH_DEPTH,
            candidates=tuple(
                Passage(doc_id=s["id"], score=s["score"], text=s["text"])
                for s in response["sources"]
            ),
        )
    )
    recorder.context(
        Context(
            items=tuple(
                ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
                for i in response["prompt_sources"]
            ),
            dropped=tuple(
                DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
                for i in response["skipped"]
            ),
            token_budget=PROMPT_WORD_BUDGET,
        )
    )
    recorder.citations(response["cited"])
    return {"answer": response["answer"], "citations": response["cited"]}
ArtifactShapeRead by
retrieval/v1query, depth, and candidates in the order your retriever returned them, each a Passage(doc_id, chunk_id, score, text)hit_rate, recall, mrr, ndcg
context/v1items that reached the model (doc_id, position, tokens, text), dropped items with a reason of top_k or token_budget, and token_budgetcitation_validity, groundedness_judge, citation_support_judge
citations/v1ids, each a doc_id or doc_id#chunk_idcitation_validity, citation_support_judge

Oloproof records positions as given and never re-ranks. A malformed artifact stops the run with exit code 2 rather than being stored. case is the case's input object, so case["question"] is the question from the dataset.

To use your own application, replace the body of support_api with a call to it (an SDK call, an HTTP request) and keep run. Point system.callable in oloproof.yaml at it as module:function.

The HTTP variant

An HTTP system cannot call the recorder, so its response carries the evidence instead, already in the three shapes above, and the configuration names where:

system:
  name: support-rag-http
  version: tutorial
  http:
    url: http://127.0.0.1:8766/answer
    output_path: result
    artifacts:
      retrieval/v1: evidence.retrieval
      context/v1: evidence.context
      citations/v1: evidence.citations

server.py serves exactly that. Start it, then run against it:

python server.py 8766
oloproof run --config oloproof.http.yaml

The case input is posted as the JSON body. output_path picks the output out of the response, and each artifacts entry records a dotted path as that kind; a missing or malformed field stops the run with exit code 2. The results are identical to the callable path below. In your own service, the evidence object is usually a debug field you enable for evaluation traffic.

What a case declares

{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}
  • expected.relevant lists the passages that answer the question. Retrieval metrics read it. A case without it, like office_hours, is excluded from them with no_relevance_labels: it leaves the denominator rather than counting as a pass or a failure.
  • expected.gold_context is the passage text itself. Path A never uses it; path B substitutes it for the retrieved context during diagnosis.

Unlabelled cases are normal in practice, since labelling relevance takes work. They still count for the answer and citation checks.

Choosing evaluators

evaluators:
  - {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
  - {type: hit_rate, k: 2}
  - {type: recall, k: 2}
  - {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4
  • contains checks the answer holds the expected text. It is the task check: did the user get the right answer. Use an exact or rubric judge instead when wording varies.
  • hit_rate and recall at k: 2 measure retrieval at the depth the application actually puts in the prompt. A retrieval metric at a depth the model never sees describes the index, not the application.
  • citation_validity checks every cited id names a passage that reached the model; require_citations: true also fails an answer that cites nothing.
  • groundedness_judge and citation_support_judge (optional) ask a model whether the answer is supported by the context. They need a provider, a model and credentials in an environment variable, and cost money per case; see Judges for what a judge must clear before it may gate.

relevant_position and context_truncated slices are not available here: they compare positions with the application's top-k, which only a staged system declares. Asking for them stops the run with slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system.

The release policy

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: answer-floor
    metric: answer_correct
    min: 0.70
  - id: retrieval-floor
    metric: hit_rate_at_2
    min: 0.80
  - id: citations-valid
    metric: citations_valid
    kind: observed_count
    max_failures: 0

A min rule passes only when the whole interval clears the floor, fails when the whole interval sits below it, and is INSUFFICIENT_EVIDENCE otherwise. An observed_count rule decides on the cases actually run, with no interval: "no invalid citation in this suite". See Gating.

Run it

oloproof run
Run run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ FAIL                  │ observed_failures_exceed_limit │

│ answer_correct  │ 73.3%    │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3%    │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 miss

How to read it:

  • Gate: BLOCK (exit 1): a rule FAILed. Exit 1 means a FAIL; exit 3 means the gate blocked without a FAIL (here it would be INSUFFICIENT_EVIDENCE); exit 0 means nothing the policy blocks on. [DECIDED/COMPLETE] is the execution state: every case ran.
  • citations-valid FAILs: one answer cited nothing, and require_citations counts that as invalid.
  • answer-floor is INSUFFICIENT_EVIDENCE, not PASS, although 73.3% is above 70%: with 15 cases the interval reaches down to 44.8%, so the evidence cannot show the floor is met.
  • hit_rate_at_2 reads 2 excluded: the two unlabelled cases. Its denominator is 13, not 15.
  • The Slices table that follows is exploratory and never gated; a slice under min_slice_support shows no interval.

Inspect the failures

The run id is on the first line of the run's output.

oloproof inspect RUN_ID --failures
4 of 15 cases failed, errored or did not finish

refund_review
  output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
  answer_correct: failed

money_back
  output: {"answer": "I could not find that in the knowledge base.", "citations": []}
  answer_correct: failed
  hit_rate_at_2: failed
  recall_at_2: failed
  citations_valid: failed

security_review
  output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
  answer_correct: failed

seat_count
  output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
  answer_correct: failed

oloproof inspect RUN_ID --case refund_review prints one case's input, expected values, output and every judgment. The recorded artifacts are in the exported bundle:

oloproof export RUN_ID

Each line of .oloproof/bundles/RUN_ID/cases.jsonl is one case's record; its artifacts field holds what was recorded. For money_back, that field reads:

{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}

Reading the four failures from the recorded evidence alone:

CaseWhat the record showsA meaningful next action
money_backRetrieval returned nothing: the question shares no word with the refund passageQuery rewriting or synonyms, measured by hit_rate_at_2
refund_review, security_reviewhit_rate_at_2 passed, yet the answer came from another passageInspect context/v1: was the relevant passage dropped for the budget?
seat_countThe right passage was retrieved, kept and cited; the answer says "five", the case expects "5"Fix the expectation or the answer format, not retrieval

That table is your reading of the record. It is an association between a failure and a stage, not a proven cause: nothing re-executed the case with the stage changed.

What diagnose does on a black box

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc

Every case is UNRESOLVED with the reason intervention_unsupported. Oloproof cannot hand a black box the gold passage in place of its own retrieval, so it does not pretend to. Controlled interventions need path B.

Make a candidate change and compare

The record says money_back failed at retrieval. The candidate change expands the question with synonyms before searching. In app.py:

EXPAND_QUERY = True

Changing the code changes the system version recorded for the run. Run again, then compare the candidate with the baseline:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml

The candidate run alone: citations-valid now PASSes, hit_rate_at_2 reads 100.0% [75.2%, 100.0%], and the gate still blocks with exit 3 because answer-floor and retrieval-floor remain INSUFFICIENT_EVIDENCE. The comparison:

Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
  excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)

The change fixed the case it targeted (one more answer, +6.7 points over 15 paired cases). The comparison still cannot establish that the candidate is not worse than the baseline by more than the margin: 15 paired cases leave an interval about 67 points wide. The planning line says how many more paired cases would decide it if the difference held. A larger suite, not a different margin, is the next action. See Comparing two runs and Comparison rules.

Path B: a staged application with diagnosis

The staged files

Path B runs the support_rag example, described on the RAG evaluation page. Copy it:

oloproof init --example support_rag my-staged-rag
cd my-staged-rag
FileWhat it is
app.pySupportRag, a class decorated with @rag_system: retrieve(input, depth), generate(input, context), count_tokens(passage)
data/corpus.jsonl, data/support.jsonlThe knowledge base, and 13 cases, each with relevant and gold_context
oloproof.yamlsystem.rag points at the class and sets depth, top_k, token_budget, index_version
release.yaml, compare.yamlThe same policies as path A

The difference from path A is who assembles the context. Here Oloproof calls retrieve, keeps the first top_k candidates, drops passages past token_budget, and passes the rest to generate. Because it holds the stages apart, it can cache them separately and re-execute generation with a different context. To adapt your own application, replace the bodies of retrieve (call your index, return Retrieval(candidates=[Passage(...)]) in your retriever's order) and generate (call your model with the passages given). Set index_version to something that changes when your index does: it is part of the retrieval's identity, and a stale value reuses cached retrievals against an index that no longer returns them.

The same configuration also allows the relevant_position and context_truncated slices, and an ndcg evaluator over the full retrieval depth.

Run the staged suite

oloproof run
Gate: BLOCK (exit 3)
│ answer-floor    │ answer_correct  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ retrieval-floor │ hit_rate_at_2   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ citations-valid │ citations_valid │ PASS                  │ observed_failures_within_limit │
│ answer_correct  │ 69.2%    │ [38.5%, 91.0%]  │ 9 / 13 observed · 0 missing · 0 excluded     │
│ hit_rate_at_2   │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ recall_at_2     │ 92.3%    │ [63.9%, 99.9%]  │ 12 / 13 observed · 0 missing · 0 excluded    │
│ ndcg_at_6       │ 0.866    │ [0.506, 0.990]  │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0%   │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded    │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 miss

The Stages line is the staged system's own cache. Exit 3: nothing FAILed, but two rules lack the evidence to PASS.

Diagnose with gold context, beside a control

oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correct
Selected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…

Two child runs are made from the failed cases: one with the case's gold_context in place of the retrieved context, and a control that re-executes them unchanged. The control is what makes the reading safe: a case that passes on a plain re-run was unstable, not diagnosed. diagnose exits 0 whatever it finds; it decides nothing about the release.

oloproof inspect DIAGNOSIS_ID
money_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2

Reading the labels

LabelWhat was observedWhat it does not establish
RETRIEVAL_MISSNo relevant passage was retrieved, and the case passed with the gold passageThat retrieval is the only thing wrong, or that a given retrieval change will fix it
RANKED_OUTA relevant passage was retrieved below top_k, and the case passed with the gold passageThat widening top-k will help other cases
CONTEXT_ASSEMBLY_LOSSA relevant passage within top-k was dropped from the context, and the case passed with the gold passageWhich budget would be enough
GENERATION_FAILUREThe case still failed with the gold passage in handThat the model, rather than the prompt or the expectation, is at fault
UNRESOLVEDNothing could be concluded: the system is not staged (intervention_unsupported), the case recovered under the control (unstable_under_control), it has no relevance labels (no_relevance_labels), or evidence is missingAnything about the case

Every label is an association between a failure and a stage under one intervention on these cases. It is not a proven cause: "Implicated" and "Candidate experiment" are the strongest words the output uses, and the counts describe the selected cases only ("no population claim"). seat_count is a good reminder: it fails with the right passage because the knowledge base says "five" and the case expects "5", which no retrieval change can fix.

Cases with and without gold passages

Only failed cases that declare expected.gold_context can be re-executed. Remove the gold passage from seat_count and money_back (and the relevance label from money_back) and the same command reports:

Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the context

The excluded cases are listed, not silently dropped. Note also what removing a relevance label does to the run itself: hit_rate_at_2 rose to 100.0% (12 / 12 observed, 1 excluded), because the one case retrieval missed is no longer measured. Unlabelled cases leave the denominator; they do not count as passes, and a metric over fewer cases can look better than the application is. Label the hard cases first.

Test a fix before making it: top-k and a reranker

Two more interventions replay the recorded retrieval with a different setting, so the retriever is not called again:

oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correct
Recovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recovered

A reranker is a function (input, candidates) -> candidates you write. Save it as rerank.py beside app.py:

"""A candidate reranker: shorter passages first, so more of them fit the token budget."""

from oloproof import Passage


def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
    return sorted(candidates, key=lambda passage: len((passage.text or "").split()))
oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correct
Recovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recovered

Neither recovers anything, which is what the gold-context labels predicted: no failure here was a passage ranked just below the cut. Each replay carries the gold-context labels forward, so the diagnoses read together.

Run the experiment the diagnosis named, and compare

Raise token_budget to 120 in oloproof.yaml, then:

oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Stages: retrieve 13 hit/0 miss; generate 7 hit/6 miss

Every retrieval was reused, because top_k and the budget are outside the retrieval's identity; only the six cases whose context changed were generated again.

answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
  answers-not-worse  answer_correct  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
  citations-not-worse  citations_valid  non-inferiority, margin 2.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
Gate: BLOCK (exit 3)

The experiment did not help: not one case changed its verdict, so the hypothesis the diagnosis offered is not supported for these cases. That is a useful result. The next experiment is smaller chunks, or the two context-assembly cases' prompts; seat_count needs its expectation fixed.

Troubleshooting

SymptomCauseFix
Configuration error: slice 'relevant_position' ... needs a staged systemA position slice on a callable or HTTP systemDrop the slice, or move to path B
Run stops with exit 2 and malformed retrieval/v1 artifactA field the schema does not allow, or more candidates than depthMap only the documented fields; set depth to at least the number returned
citations_valid reads 0 / 0 observed · 15 missing and its rule is INSUFFICIENT_EVIDENCE with no_observationsThe adapter did not record citations/v1 (or context/v1); each such case is missing, not passedRecord both on every path through the adapter, including "no answer"; oloproof inspect RUN_ID --failures shows the error per case
A retrieval metric shows many excludedCases without expected.relevantLabel them, or accept the smaller denominator knowingly
diagnose says UNRESOLVED ... not stagedPath AExpected; use path B for interventions
diagnose refuses with an intervention must re-execute the same systemThe code or config changed since the runDiagnose a run of the current version, or restore the version that ran
Diagnosis selects fewer cases than failedFailed cases without expected.gold_contextAdd the gold passages; the excluded cases are named in the output
Retrievals are reused after the index changedindex_version unchangedChange index_version when the index changes

Limitations

  • Oloproof calls your application; it does not host, sandbox or reset it. Its index, caches and any state it keeps are yours.
  • On a black box, interventions are not available: diagnose labels every case UNRESOLVED and re-executes nothing.
  • The interventions are gold context, top-k and a reranker. There is no chunking, embedding or prompt intervention.
  • Diagnosis labels describe the selected failed cases under one intervention beside a control. They associate a failure with a stage; they do not prove a cause, and they make no claim about cases not selected.
  • Retrieval metrics need relevance labels, and the diagnosis needs gold passages; Oloproof does not create either.
  • The deterministic examples stand in for a real retriever and model. A live model in generate or a judge evaluator calls a provider, needs credentials and costs money per case.
  • What works where, SDK against YAML against browser, is on What works today.