Skip to content

For RAG and retrieval teams

Know what your index change did.

Retrieval changes move quality in ways a chunking diff cannot show. Oloproof runs your RAG system as two stages it can see, and measures retrieval, citations, groundedness and answer correctness against the same cases, before and after.

Measure

Retrieval-aware evaluators

Hit rate, recall, MRR and nDCG at the cutoff you choose, read from the recorded candidates; citation validity against the context the model was given; groundedness and citation-support judges that read it. A cutoff deeper than the retriever returned is refused, not scored as a miss.

Isolate

One stage at a time

Each stage is cached on its own, so a prompt change never searches again. Diagnosis re-runs failed cases with the gold passage beside a control and labels each a retrieval miss, ranked out, dropped from context, or a generation failure.

Decide

Gate on the interval

Candidate and baseline are paired case by case against margins you declare. A rule passes only when its interval clears the margin; an interval too wide to tell blocks the release as insufficient evidence, with the evidence kept beside the decision.

A real result

A live RAG docs assistant · rerank on against rerank off · 58 questions · our own run, 22 September 2026

Right page at rank 1

+6.2 points

[−36.3, +45.5] · 48 paired · undecided

Right page in the top 3

+3.7 points

[−73.9, +76.2] · 27 paired · undecided

Answer correct

+6.5 points

[−39.4, +48.4] · 46 paired · undecided

Read off the point estimates, the reranker added 12.7 points of right page at rank 1 and would have shipped. Paired on the questions both arms answered, every interval crosses zero: this panel cannot say whether the reranker helps or hurts, and Oloproof says so instead of picking a winner. Read the field report.