For RAG and retrieval teams
Know what your index change did.
Retrieval changes move quality in ways a chunking diff cannot show. Oloproof runs your RAG system as two stages it can see, and measures retrieval, citations, groundedness and answer correctness against the same cases, before and after.
Measure
Retrieval-aware evaluators
Hit rate, recall, MRR and nDCG at the cutoff you choose, read from the recorded candidates; citation validity against the context the model was given; groundedness and citation-support judges that read it. A cutoff deeper than the retriever returned is refused, not scored as a miss.
Isolate
One stage at a time
Each stage is cached on its own, so a prompt change never searches again. Diagnosis re-runs failed cases with the gold passage beside a control and labels each a retrieval miss, ranked out, dropped from context, or a generation failure.
Decide
Gate on the interval
Candidate and baseline are paired case by case against margins you declare. A rule passes only when its interval clears the margin; an interval too wide to tell blocks the release as insufficient evidence, with the evidence kept beside the decision.
A real result
Right page at rank 1
+6.2 points
[−36.3, +45.5] · 48 paired · undecided
Right page in the top 3
+3.7 points
[−73.9, +76.2] · 27 paired · undecided
Answer correct
+6.5 points
[−39.4, +48.4] · 46 paired · undecided
Read off the point estimates, the reranker added 12.7 points of right page at rank 1 and would have shipped. Paired on the questions both arms answered, every interval crosses zero: this panel cannot say whether the reranker helps or hurts, and Oloproof says so instead of picking a winner. Read the field report.