Skip to content

Field report · our own run

The reranker looked like a win. The evidence could not tell.

We ran Oloproof against a live retrieval-augmented docs assistant that we do not operate: its own retriever and index, real generation, an LLM judge, and its owners' 58-question golden panel, unchanged. No customer is involved. This is our run, reported as we audited it.

58
questions in the owners' golden panel
0 of 76
feature-flag comparisons settled
1 in 9
cases changing verdict between identical runs
15 of 23
failed answers recovered with the gold passage

Rerank on against rerank off, 22 September 2026

The reranker comparison: each arm's rate with its 95% interval and observed count, and the paired difference
MetricRerank onRerank offPaired difference, on minus off
Right page at rank 10.745 [0.519, 0.875] · 51 of 580.618 [0.449, 0.760] · 55 of 58+6.2 points [−36.3, +45.5] · 48 paired
Right page in the top 30.742 [0.270, 0.939] · 31 of 580.683 [0.350, 0.875] · 41 of 58+3.7 points [−73.9, +76.2] · 27 paired
Answer correct (LLM judge)0.760 [0.519, 0.888] · 50 of 580.685 [0.501, 0.819] · 54 of 58+6.5 points [−39.4, +48.4] · 46 paired

Before

The assistant already had a capable harness: a golden panel, an LLM judge and a release gate. The gate compared a point estimate with a target, at every panel size. Read that way, the reranker added 12.7 points of right-page-at-rank-1 and would have shipped. And in the faithfulness score that gated releases, an answer the judge failed to grade counted as a total failure.

After

Oloproof paired the two arms on the questions both had answered, recorded 4 server errors and 2 unparseable verdicts as missing, and widened each interval to bound them. All three differences cross zero, widely: this panel cannot say whether the reranker helps or hurts.

Re-running the 23 failed answers with the gold passage in place of the retrieved one recovered 15, beside a control that re-ran them unchanged. The diagnosis labelled 5 retrieval misses, 7 generation failures and left 11 unresolved, and named retrieval depth and the index configuration as the next experiments.

Measured three times over, about one case in nine changed its verdict with nothing changed, and 13.8% of calls returned a server error. Declaring those errors transient recovered 18 of 22 missing cases and narrowed the interval by 39%. A sweep of the assistant's nineteen feature flags then gave 76 paired comparisons, none of them settled, and about 40 more questions would settle the three most promising.

“A suite of 58 cases, one in nine of which answers differently each time it is asked, does not have four decimal places of resolution.”
From Oloproof's report of the run, 23 September 2026

At a glance

  • Run byOloproof, on its own initiative
  • System evaluatedA live RAG docs assistant we do not operate
  • Panel58 questions, the owners' own, unchanged
  • JudgeAn LLM judge, not yet checked against human labels
  • Run on22 and 23 September 2026
  • Follow-up runs19 feature flags; 24 conversations

What it does not show

One system, one corpus, and a panel too small to settle its own question. The judge behind every answer score has not been validated against people, and the diagnosis says so. The panel is single-turn: a follow-up run of 24 conversations found two reproducible failures that single-turn panels cannot produce, including an assistant denying what it had said one turn earlier. It is not a benchmark, and it is not a customer.