Take a golden panel of 58 cases and a live documentation assistant that answers from a real search index, and compare it with its reranker on and with it off. Read off the point estimates, the reranker raises the share of answers whose top-ranked page is the right one from 61.8% to 74.5%. That reads like progress, and in a status update it usually gets reported as a 12.7-point gain.
Run the same arm twice with nothing changed and answer correctness came back 71.2%, then 76.0%. Measured again with three replicates per case, about one observed case in nine changed its verdict between identical measurements: 6 of 54 on correctness, 7 of 57 on the top-ranked page.
The reason is not that the measurement is broken. It is that a difference on 58 cases, some of which could not be measured at all, is well inside the noise the sample can produce on its own. The interval, not the point estimate, tells you whether you learned anything.
What to report
Three numbers, in this order: the delta, its interval, and the number of cases behind it. Anything less and a reader cannot judge the claim. Here is that run, reranker on against reranker off, paired case by case, in points:
| Paired cases | Metric | Delta | 95% interval | Reads as |
|---|---|---|---|---|
| 48 | Right page at rank 1 | +6.2 | −36.3 … +45.5 | Inconclusive |
| 46 | Answer correct | +6.5 | −39.4 … +48.4 | Inconclusive |
| 27 | Right page in the top 3 | +3.7 | −73.9 … +76.2 | Inconclusive |
Every interval crosses zero, widely. The point estimates all say the reranker helps; the intervals say this panel cannot tell whether it helps or hurts. That is a finding in its own right: the panel is too small, and too lossy, to settle the question it was being used to settle.
Missing is not zero
Four of the 58 calls returned HTTP 500 from the system under test, and two judge calls returned no verdict that could be parsed. A harness that averages an ungraded case as zero records those six as the assistant failing: six manufactured failures out of 58. Oloproof records them as missing and widens the interval to bound whatever they might have been.
That is what the third row is pricing. Only 27 of the 58 cases could be measured at a cutoff of three ranked pages, and the interval accounts for the other 31 instead of hiding them.
Small panels pass gates
The same rule holds at the small end. A smoke run of three cases scored 1.000, with a 95% interval of 0.292 to 1.000. A gate that compares the point estimate with a target ships it. A gate that reads the interval says what is true: three cases cannot tell a perfect system from one that is right three times in ten.
Narrowing the interval
More cases narrow an interval, and so does losing fewer of them. On the same panel, declaring the system's HTTP 500s transient, so that the runner retried them, recovered 18 of 22 missing cases and narrowed the interval by 39%, minutes later, with nothing else changed.
A delta whose interval clears the threshold supports a decision, whatever its size, and that is the one a gate should act on. Until then, report the delta, its interval and the cases behind it, and say the answer is not in yet.
Want this computed for you on every change?
Contact us