Field report · our own run
The reranker looked like a win. The evidence could not tell.
We ran Oloproof against a live retrieval-augmented docs assistant that we do not operate: its own retriever and index, real generation, an LLM judge, and its owners' 58-question golden panel, unchanged. No customer is involved. This is our run, reported as we audited it.
- 58
- questions in the owners' golden panel
- 0 of 76
- feature-flag comparisons settled
- 1 in 9
- cases changing verdict between identical runs
- 15 of 23
- failed answers recovered with the gold passage
Rerank on against rerank off, 22 September 2026
| Metric | Rerank on | Rerank off | Paired difference, on minus off |
|---|---|---|---|
| Right page at rank 1 | 0.745 [0.519, 0.875] · 51 of 58 | 0.618 [0.449, 0.760] · 55 of 58 | +6.2 points [−36.3, +45.5] · 48 paired |
| Right page in the top 3 | 0.742 [0.270, 0.939] · 31 of 58 | 0.683 [0.350, 0.875] · 41 of 58 | +3.7 points [−73.9, +76.2] · 27 paired |
| Answer correct (LLM judge) | 0.760 [0.519, 0.888] · 50 of 58 | 0.685 [0.501, 0.819] · 54 of 58 | +6.5 points [−39.4, +48.4] · 46 paired |
Before
The assistant already had a capable harness: a golden panel, an LLM judge and a release gate. The gate compared a point estimate with a target, at every panel size. Read that way, the reranker added 12.7 points of right-page-at-rank-1 and would have shipped. And in the faithfulness score that gated releases, an answer the judge failed to grade counted as a total failure.
After
Oloproof paired the two arms on the questions both had answered, recorded 4 server errors and 2 unparseable verdicts as missing, and widened each interval to bound them. All three differences cross zero, widely: this panel cannot say whether the reranker helps or hurts.
Re-running the 23 failed answers with the gold passage in place of the retrieved one recovered 15, beside a control that re-ran them unchanged. The diagnosis labelled 5 retrieval misses, 7 generation failures and left 11 unresolved, and named retrieval depth and the index configuration as the next experiments.
Measured three times over, about one case in nine changed its verdict with nothing changed, and 13.8% of calls returned a server error. Declaring those errors transient recovered 18 of 22 missing cases and narrowed the interval by 39%. A sweep of the assistant's nineteen feature flags then gave 76 paired comparisons, none of them settled, and about 40 more questions would settle the three most promising.
“A suite of 58 cases, one in nine of which answers differently each time it is asked, does not have four decimal places of resolution.”
At a glance
- Run byOloproof, on its own initiative
- System evaluatedA live RAG docs assistant we do not operate
- Panel58 questions, the owners' own, unchanged
- JudgeAn LLM judge, not yet checked against human labels
- Run on22 and 23 September 2026
- Follow-up runs19 feature flags; 24 conversations
What it does not show
One system, one corpus, and a panel too small to settle its own question. The judge behind every answer score has not been validated against people, and the diagnosis says so. The panel is single-turn: a follow-up run of 24 conversations found two reproducible failures that single-turn panels cannot produce, including an assistant denying what it had said one turn earlier. It is not a benchmark, and it is not a customer.