Guides
Tutorial: evaluate a RAG application
A runnable walkthrough for a retrieval-augmented application, in two paths: an existing black-box application whose retrieval, context and citations you record from the outside, and a staged application Oloproof runs stage by stage so oloproof diagnose can re-execute failed cases under controlled changes. Both run locally with no provider credentials.
The concepts behind each step (stages, relevance labels, gold context, the four failure labels) are on the RAG evaluation page; the terms case, evaluator, metric, interval and gate are in Core concepts. This page is the hands-on route through them.
Which path is yours
| Your application | Path | What you get | What you do not get |
|---|---|---|---|
| One call in, one answer out (a service, an HTTP endpoint, a framework chain you do not want to split) | A, black box | Retrieval metrics, citation checks, grounding judges, gating, comparison | Controlled interventions: diagnose re-executes nothing |
| Retrieval and generation you can call separately | B, staged | Everything in A, per-stage caching, and diagnose with gold context, top-k and a reranker beside a control | Interventions other than those three |
Start with A if you are unsure. It needs no change to the application, and moving to B later keeps the dataset, evaluators and policy.
Prerequisites
- Python 3.11 or later, and Oloproof installed (pip install oloproof).
- The example projects, which ship with the package: blackbox_rag for path A and support_rag for path B. Copy one into a new directory and work there:
oloproof init --example blackbox_rag my-rag
cd my-ragEvery command below runs from inside the copied directory. Runs, judgments and diagnoses are stored in .oloproof/ there.
Path A: an existing application as a black box
The files
| File | What it is |
|---|---|
| app.py | support_api(question), standing in for your application, and run(case), the adapter |
| server.py | The same application over HTTP, for the HTTP variant below |
| data/corpus.jsonl | The 14-passage knowledge base the application searches |
| data/support.jsonl | 15 cases: 13 with relevance labels and gold passages, 2 without either |
| oloproof.yaml | The suite: dataset, system, evaluators, slices |
| oloproof.http.yaml | The same suite against the HTTP server |
| release.yaml | The release policy for a single run |
| compare.yaml | The policy for comparing a candidate run with a baseline |
What the application returns
support_api behaves like an application you already have: it searches, builds a prompt from the best sources that fit a word budget, answers and cites. Its response already carries what it did:
{
"answer": "Team plans include five seats.",
"cited": ["kb-03"],
"sources": [{"id": "kb-03", "score": 3.0, "text": "Team plans include five seats. ..."}],
"prompt_sources": [{"id": "kb-03", "score": 3.0, "text": "...", "rank": 1, "tokens": 17}],
"skipped": [{"id": "kb-05", "rank": 3, "why": "top_k"}]
}Your application's field names will differ. What matters is that it can tell you, per question, the ranked sources it retrieved, the ones that reached the model, and the ones it cited. If it cannot, add those to its response or its logs first: Oloproof measures what is recorded and never infers retrieval from an answer.
The adapter
run calls the application unchanged and maps the response onto three typed artifacts, the records the retrieval and citation evaluators read:
@system(
name="support-rag-blackbox",
version="tutorial",
records=("retrieval/v1", "context/v1", "citations/v1"),
)
def run(case):
response = support_api(str(case["question"]))
recorder = current_case()
recorder.retrieval(
Retrieval(
query=case["question"],
depth=SEARCH_DEPTH,
candidates=tuple(
Passage(doc_id=s["id"], score=s["score"], text=s["text"])
for s in response["sources"]
),
)
)
recorder.context(
Context(
items=tuple(
ContextItem(doc_id=i["id"], position=i["rank"], tokens=i["tokens"], text=i["text"])
for i in response["prompt_sources"]
),
dropped=tuple(
DroppedItem(doc_id=i["id"], position=i["rank"], reason=i["why"])
for i in response["skipped"]
),
token_budget=PROMPT_WORD_BUDGET,
)
)
recorder.citations(response["cited"])
return {"answer": response["answer"], "citations": response["cited"]}| Artifact | Shape | Read by |
|---|---|---|
| retrieval/v1 | query, depth, and candidates in the order your retriever returned them, each a Passage(doc_id, chunk_id, score, text) | hit_rate, recall, mrr, ndcg |
| context/v1 | items that reached the model (doc_id, position, tokens, text), dropped items with a reason of top_k or token_budget, and token_budget | citation_validity, groundedness_judge, citation_support_judge |
| citations/v1 | ids, each a doc_id or doc_id#chunk_id | citation_validity, citation_support_judge |
Oloproof records positions as given and never re-ranks. A malformed artifact stops the run with exit code 2 rather than being stored. case is the case's input object, so case["question"] is the question from the dataset.
To use your own application, replace the body of support_api with a call to it (an SDK call, an HTTP request) and keep run. Point system.callable in oloproof.yaml at it as module:function.
The HTTP variant
An HTTP system cannot call the recorder, so its response carries the evidence instead, already in the three shapes above, and the configuration names where:
system:
name: support-rag-http
version: tutorial
http:
url: http://127.0.0.1:8766/answer
output_path: result
artifacts:
retrieval/v1: evidence.retrieval
context/v1: evidence.context
citations/v1: evidence.citationsserver.py serves exactly that. Start it, then run against it:
python server.py 8766
oloproof run --config oloproof.http.yamlThe case input is posted as the JSON body. output_path picks the output out of the response, and each artifacts entry records a dotted path as that kind; a missing or malformed field stops the run with exit code 2. The results are identical to the callable path below. In your own service, the evidence object is usually a debug field you enable for evaluation traffic.
What a case declares
{"id":"seat_count","input":{"question":"How many seats does a team plan include?"},"expected":{"answer":"5 seats","relevant":[{"doc_id":"kb-03"}],"gold_context":[{"doc_id":"kb-03","text":"Team plans include five seats. ..."}]},"metadata":{"topic":"billing"}}
{"id":"office_hours","input":{"question":"What are the support office hours?"},"expected":{"answer":"09:00"},"metadata":{"topic":"account"}}- expected.relevant lists the passages that answer the question. Retrieval metrics read it. A case without it, like office_hours, is excluded from them with no_relevance_labels: it leaves the denominator rather than counting as a pass or a failure.
- expected.gold_context is the passage text itself. Path A never uses it; path B substitutes it for the retrieved context during diagnosis.
Unlabelled cases are normal in practice, since labelling relevance takes work. They still count for the answer and citation checks.
Choosing evaluators
evaluators:
- {type: contains, criterion: answer_correct, field: answer, expected_field: answer}
- {type: hit_rate, k: 2}
- {type: recall, k: 2}
- {type: citation_validity, require_citations: true}
slices: [metadata.topic]
min_slice_support: 4- contains checks the answer holds the expected text. It is the task check: did the user get the right answer. Use an exact or rubric judge instead when wording varies.
- hit_rate and recall at k: 2 measure retrieval at the depth the application actually puts in the prompt. A retrieval metric at a depth the model never sees describes the index, not the application.
- citation_validity checks every cited id names a passage that reached the model; require_citations: true also fails an answer that cites nothing.
- groundedness_judge and citation_support_judge (optional) ask a model whether the answer is supported by the context. They need a provider, a model and credentials in an environment variable, and cost money per case; see Judges for what a judge must clear before it may gate.
relevant_position and context_truncated slices are not available here: they compare positions with the application's top-k, which only a staged system declares. Asking for them stops the run with slice 'relevant_position' compares relevant positions with top_k, so it needs a staged system.
The release policy
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: answer-floor
metric: answer_correct
min: 0.70
- id: retrieval-floor
metric: hit_rate_at_2
min: 0.80
- id: citations-valid
metric: citations_valid
kind: observed_count
max_failures: 0A min rule passes only when the whole interval clears the floor, fails when the whole interval sits below it, and is INSUFFICIENT_EVIDENCE otherwise. An observed_count rule decides on the cases actually run, with no interval: "no invalid citation in this suite". See Gating.
Run it
oloproof runRun run_01M4FCBPE0G550CKCVGXCNEM2P [DECIDED/COMPLETE]
Gate: BLOCK (exit 1)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ FAIL │ observed_failures_exceed_limit │
│ answer_correct │ 73.3% │ [44.8%, 92.3%] │ 11 / 15 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 2 excluded │
│ citations_valid │ 93.3% │ [68.0%, 99.9%] │ 14 / 15 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/15 miss; judgment 0 hit/56 missHow to read it:
- Gate: BLOCK (exit 1): a rule FAILed. Exit 1 means a FAIL; exit 3 means the gate blocked without a FAIL (here it would be INSUFFICIENT_EVIDENCE); exit 0 means nothing the policy blocks on. [DECIDED/COMPLETE] is the execution state: every case ran.
- citations-valid FAILs: one answer cited nothing, and require_citations counts that as invalid.
- answer-floor is INSUFFICIENT_EVIDENCE, not PASS, although 73.3% is above 70%: with 15 cases the interval reaches down to 44.8%, so the evidence cannot show the floor is met.
- hit_rate_at_2 reads 2 excluded: the two unlabelled cases. Its denominator is 13, not 15.
- The Slices table that follows is exploratory and never gated; a slice under min_slice_support shows no interval.
Inspect the failures
The run id is on the first line of the run's output.
oloproof inspect RUN_ID --failures4 of 15 cases failed, errored or did not finish
refund_review
output: {"answer": "Every refund request on an annual plan is logged in the audit trail, and the same request is listed again on the day it was reviewed and approved."…
answer_correct: failed
money_back
output: {"answer": "I could not find that in the knowledge base.", "citations": []}
answer_correct: failed
hit_rate_at_2: failed
recall_at_2: failed
citations_valid: failed
security_review
output: {"answer": "Security reviews during Enterprise onboarding include an access review and a written summary for the customer, and every review is scheduled with t…
answer_correct: failed
seat_count
output: {"answer": "Team plans include five seats.", "citations": ["kb-03"]}
answer_correct: failedoloproof inspect RUN_ID --case refund_review prints one case's input, expected values, output and every judgment. The recorded artifacts are in the exported bundle:
oloproof export RUN_IDEach line of .oloproof/bundles/RUN_ID/cases.jsonl is one case's record; its artifacts field holds what was recorded. For money_back, that field reads:
{"retrieval/v1": [{"candidates": [], "depth": 6, "query": "Where do I claim money back on a yearly subscription?"}], "context/v1": [{"dropped": [], "items": [], "source": "retrieval", "token_budget": 40}], "citations/v1": [{"ids": []}]}Reading the four failures from the recorded evidence alone:
| Case | What the record shows | A meaningful next action |
|---|---|---|
| money_back | Retrieval returned nothing: the question shares no word with the refund passage | Query rewriting or synonyms, measured by hit_rate_at_2 |
| refund_review, security_review | hit_rate_at_2 passed, yet the answer came from another passage | Inspect context/v1: was the relevant passage dropped for the budget? |
| seat_count | The right passage was retrieved, kept and cited; the answer says "five", the case expects "5" | Fix the expectation or the answer format, not retrieval |
That table is your reading of the record. It is an association between a failure and a stage, not a proven cause: nothing re-executed the case with the stage changed.
What diagnose does on a black box
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
UNRESOLVED: 4 of 4, the system is not staged, so no case was re-executed
Diagnosis sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bc
Cases: oloproof inspect sha256:8809e100ab2ec1dad8ffeacc136cd510300081a0edf01ec11f881c543ed104bcEvery case is UNRESOLVED with the reason intervention_unsupported. Oloproof cannot hand a black box the gold passage in place of its own retrieval, so it does not pretend to. Controlled interventions need path B.
Make a candidate change and compare
The record says money_back failed at retrieval. The candidate change expands the question with synonyms before searching. In app.py:
EXPAND_QUERY = TrueChanging the code changes the system version recorded for the run. Run again, then compare the candidate with the baseline:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlThe candidate run alone: citations-valid now PASSes, hit_rate_at_2 reads 100.0% [75.2%, 100.0%], and the gate still blocks with exit 3 because answer-floor and retrieval-floor remain INSUFFICIENT_EVIDENCE. The comparison:
Comparison sha256:2feb024c… of run_01M4FCCJYCVVYA8NB4XZDV6YMB against run_01M4FCCHVBG9WHDP7G5HFX7RDT · 15 paired cases
answer_correct: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
hit_rate_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
recall_at_2: +7.7 points [-29.8, +45.5] · 13 paired · 0 missing · 2 excluded
excluded 2: no_relevance_labels
citations_valid: +6.7 points [-26.5, +40.8] · 15 paired · 0 missing · 0 excluded
20 exploratory slice differences not shown; add --slices to list them
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 38 more paired cases would decide it, if the difference holds (53 in total at 7% discordance)
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 68 more paired cases would decide it, if the difference holds (83 in total at 7% discordance)
Gate: BLOCK (exit 3)The change fixed the case it targeted (one more answer, +6.7 points over 15 paired cases). The comparison still cannot establish that the candidate is not worse than the baseline by more than the margin: 15 paired cases leave an interval about 67 points wide. The planning line says how many more paired cases would decide it if the difference held. A larger suite, not a different margin, is the next action. See Comparing two runs and Comparison rules.
Path B: a staged application with diagnosis
The staged files
Path B runs the support_rag example, described on the RAG evaluation page. Copy it:
oloproof init --example support_rag my-staged-rag
cd my-staged-rag| File | What it is |
|---|---|
| app.py | SupportRag, a class decorated with @rag_system: retrieve(input, depth), generate(input, context), count_tokens(passage) |
| data/corpus.jsonl, data/support.jsonl | The knowledge base, and 13 cases, each with relevant and gold_context |
| oloproof.yaml | system.rag points at the class and sets depth, top_k, token_budget, index_version |
| release.yaml, compare.yaml | The same policies as path A |
The difference from path A is who assembles the context. Here Oloproof calls retrieve, keeps the first top_k candidates, drops passages past token_budget, and passes the rest to generate. Because it holds the stages apart, it can cache them separately and re-execute generation with a different context. To adapt your own application, replace the bodies of retrieve (call your index, return Retrieval(candidates=[Passage(...)]) in your retriever's order) and generate (call your model with the passages given). Set index_version to something that changes when your index does: it is part of the retrieval's identity, and a stale value reuses cached retrievals against an index that no longer returns them.
The same configuration also allows the relevant_position and context_truncated slices, and an ndcg evaluator over the full retrieval depth.
Run the staged suite
oloproof runGate: BLOCK (exit 3)
│ answer-floor │ answer_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ retrieval-floor │ hit_rate_at_2 │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ citations-valid │ citations_valid │ PASS │ observed_failures_within_limit │
│ answer_correct │ 69.2% │ [38.5%, 91.0%] │ 9 / 13 observed · 0 missing · 0 excluded │
│ hit_rate_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ recall_at_2 │ 92.3% │ [63.9%, 99.9%] │ 12 / 13 observed · 0 missing · 0 excluded │
│ ndcg_at_6 │ 0.866 │ [0.506, 0.990] │ mean of 13 observed · 0 missing · 0 excluded │
│ citations_valid │ 100.0% │ [75.2%, 100.0%] │ 13 / 13 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/13 miss; judgment 0 hit/65 miss
Stages: retrieve 0 hit/13 miss; generate 0 hit/13 missThe Stages line is the staged system's own cache. Exit 3: nothing FAILed, but two rules lack the evidence to PASS.
Diagnose with gold context, beside a control
oloproof diagnose RUN_ID --intervention gold-context --criterion answer_correctSelected: 4 failed cases with gold context (observed; no population claim)
Control: 0 of 4 passed when re-executed without the intervention
Recovered under gold context: 3 of 4
RETRIEVAL_MISS: 1 of 4, recovered; no relevant evidence was retrieved
CONTEXT_ASSEMBLY_LOSS: 2 of 4, recovered; relevant evidence within top-k was left out of the context
GENERATION_FAILURE: 1 of 4, still failed with the gold context
Implicated: context budget, in 2 of the 3 recovered failures.
Candidate experiment: a larger token budget. This is a hypothesis to test, not an established cause.
Candidate experiment: smaller chunks. This is a hypothesis to test, not an established cause.
Diagnosis sha256:50a6124f…
Child runs: gold context run_…, control run_…
Cases: oloproof inspect sha256:50a6124f…Two child runs are made from the failed cases: one with the case's gold_context in place of the retrieved context, and a control that re-executes them unchanged. The control is what makes the reading safe: a case that passes on a plain re-run was unstable, not diagnosed. diagnose exits 0 whatever it finds; it decides nothing about the release.
oloproof inspect DIAGNOSIS_IDmoney_back: RETRIEVAL_MISS, relevant_not_retrieved, strength intervention_recovery, best relevant position none
refund_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2
seat_count: GENERATION_FAILURE, fails_with_gold_context, strength intervention_non_recovery, best relevant position 1
security_review: CONTEXT_ASSEMBLY_LOSS, relevant_dropped_from_context, strength intervention_recovery, best relevant position 2Reading the labels
| Label | What was observed | What it does not establish |
|---|---|---|
| RETRIEVAL_MISS | No relevant passage was retrieved, and the case passed with the gold passage | That retrieval is the only thing wrong, or that a given retrieval change will fix it |
| RANKED_OUT | A relevant passage was retrieved below top_k, and the case passed with the gold passage | That widening top-k will help other cases |
| CONTEXT_ASSEMBLY_LOSS | A relevant passage within top-k was dropped from the context, and the case passed with the gold passage | Which budget would be enough |
| GENERATION_FAILURE | The case still failed with the gold passage in hand | That the model, rather than the prompt or the expectation, is at fault |
| UNRESOLVED | Nothing could be concluded: the system is not staged (intervention_unsupported), the case recovered under the control (unstable_under_control), it has no relevance labels (no_relevance_labels), or evidence is missing | Anything about the case |
Every label is an association between a failure and a stage under one intervention on these cases. It is not a proven cause: "Implicated" and "Candidate experiment" are the strongest words the output uses, and the counts describe the selected cases only ("no population claim"). seat_count is a good reminder: it fails with the right passage because the knowledge base says "five" and the case expects "5", which no retrieval change can fix.
Cases with and without gold passages
Only failed cases that declare expected.gold_context can be re-executed. Remove the gold passage from seat_count and money_back (and the relevance label from money_back) and the same command reports:
Selected: 2 failed cases with gold context (observed; no population claim)
Excluded: 2 failed cases, no_gold_context - declare the passages that would have answered the case in its `expected.gold_context`, as a list of `{doc_id, text}` objects; an intervention needs them to tell a retrieval failure from a generation one
Control: 0 of 2 passed when re-executed without the intervention
Recovered under gold context: 2 of 2
CONTEXT_ASSEMBLY_LOSS: 2 of 2, recovered; relevant evidence within top-k was left out of the contextThe excluded cases are listed, not silently dropped. Note also what removing a relevance label does to the run itself: hit_rate_at_2 rose to 100.0% (12 / 12 observed, 1 excluded), because the one case retrieval missed is no longer measured. Unlabelled cases leave the denominator; they do not count as passes, and a metric over fewer cases can look better than the application is. Label the hard cases first.
Test a fix before making it: top-k and a reranker
Two more interventions replay the recorded retrieval with a different setting, so the retriever is not called again:
oloproof diagnose RUN_ID --intervention top-k --top-k 4 --criterion answer_correctRecovered under top-k 4: 0 of 4
Confirmed under top-k 4: 0 of 0 RANKED_OUT cases also recovered
Labels from gold context (diagnosis sha256:50a6124f…): 3 of 4 recoveredA reranker is a function (input, candidates) -> candidates you write. Save it as rerank.py beside app.py:
"""A candidate reranker: shorter passages first, so more of them fit the token budget."""
from oloproof import Passage
def shortest_first(input: dict, candidates: list[Passage]) -> list[Passage]:
return sorted(candidates, key=lambda passage: len((passage.text or "").split()))oloproof diagnose RUN_ID --intervention reranker --reranker rerank:shortest_first --criterion answer_correctRecovered under reranker rerank:shortest_first: 0 of 4
Confirmed under reranker rerank:shortest_first: 0 of 0 RANKED_OUT cases also recoveredNeither recovers anything, which is what the gold-context labels predicted: no failure here was a passage ranked just below the cut. Each replay carries the gold-context labels forward, so the diagnoses read together.
Run the experiment the diagnosis named, and compare
Raise token_budget to 120 in oloproof.yaml, then:
oloproof run
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlStages: retrieve 13 hit/0 miss; generate 7 hit/6 missEvery retrieval was reused, because top_k and the budget are outside the retrieval's identity; only the six cases whose context changed were generated again.
answer_correct: +0.0 points [-33.6, +33.6] · 13 paired · 0 missing · 0 excluded
Decisions
answers-not-worse answer_correct non-inferiority, margin 5.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
citations-not-worse citations_valid non-inferiority, margin 2.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
Gate: BLOCK (exit 3)The experiment did not help: not one case changed its verdict, so the hypothesis the diagnosis offered is not supported for these cases. That is a useful result. The next experiment is smaller chunks, or the two context-assembly cases' prompts; seat_count needs its expectation fixed.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| Configuration error: slice 'relevant_position' ... needs a staged system | A position slice on a callable or HTTP system | Drop the slice, or move to path B |
| Run stops with exit 2 and malformed retrieval/v1 artifact | A field the schema does not allow, or more candidates than depth | Map only the documented fields; set depth to at least the number returned |
| citations_valid reads 0 / 0 observed · 15 missing and its rule is INSUFFICIENT_EVIDENCE with no_observations | The adapter did not record citations/v1 (or context/v1); each such case is missing, not passed | Record both on every path through the adapter, including "no answer"; oloproof inspect RUN_ID --failures shows the error per case |
| A retrieval metric shows many excluded | Cases without expected.relevant | Label them, or accept the smaller denominator knowingly |
| diagnose says UNRESOLVED ... not staged | Path A | Expected; use path B for interventions |
| diagnose refuses with an intervention must re-execute the same system | The code or config changed since the run | Diagnose a run of the current version, or restore the version that ran |
| Diagnosis selects fewer cases than failed | Failed cases without expected.gold_context | Add the gold passages; the excluded cases are named in the output |
| Retrievals are reused after the index changed | index_version unchanged | Change index_version when the index changes |
Limitations
- Oloproof calls your application; it does not host, sandbox or reset it. Its index, caches and any state it keeps are yours.
- On a black box, interventions are not available: diagnose labels every case UNRESOLVED and re-executes nothing.
- The interventions are gold context, top-k and a reranker. There is no chunking, embedding or prompt intervention.
- Diagnosis labels describe the selected failed cases under one intervention beside a control. They associate a failure with a stage; they do not prove a cause, and they make no claim about cases not selected.
- Retrieval metrics need relevance labels, and the diagnosis needs gold passages; Oloproof does not create either.
- The deterministic examples stand in for a real retriever and model. A live model in generate or a judge evaluator calls a provider, needs credentials and costs money per case.
- What works where, SDK against YAML against browser, is on What works today.