Skip to content
← All writing

Your eval scores are lying about sample size

A single average is the most common way teams fool themselves about an AI change. Here is what we report instead, and why the interval belongs in front of the number.

Take a golden panel of 58 cases and a live documentation assistant that answers from a real search index, and compare it with its reranker on and with it off. Read off the point estimates, the reranker raises the share of answers whose top-ranked page is the right one from 61.8% to 74.5%. That reads like progress, and in a status update it usually gets reported as a 12.7-point gain.

Run the same arm twice with nothing changed and answer correctness came back 71.2%, then 76.0%. Measured again with three replicates per case, about one observed case in nine changed its verdict between identical measurements: 6 of 54 on correctness, 7 of 57 on the top-ranked page.

The reason is not that the measurement is broken. It is that a difference on 58 cases, some of which could not be measured at all, is well inside the noise the sample can produce on its own. The interval, not the point estimate, tells you whether you learned anything.

What to report

Three numbers, in this order: the delta, its interval, and the number of cases behind it. Anything less and a reader cannot judge the claim. Here is that run, reranker on against reranker off, paired case by case, in points:

Paired differences, reranker on minus reranker off, on a 58-case panel
Paired casesMetricDelta95% intervalReads as
48Right page at rank 1+6.2−36.3 … +45.5Inconclusive
46Answer correct+6.5−39.4 … +48.4Inconclusive
27Right page in the top 3+3.7−73.9 … +76.2Inconclusive

Every interval crosses zero, widely. The point estimates all say the reranker helps; the intervals say this panel cannot tell whether it helps or hurts. That is a finding in its own right: the panel is too small, and too lossy, to settle the question it was being used to settle.

Missing is not zero

Four of the 58 calls returned HTTP 500 from the system under test, and two judge calls returned no verdict that could be parsed. A harness that averages an ungraded case as zero records those six as the assistant failing: six manufactured failures out of 58. Oloproof records them as missing and widens the interval to bound whatever they might have been.

That is what the third row is pricing. Only 27 of the 58 cases could be measured at a cutoff of three ranked pages, and the interval accounts for the other 31 instead of hiding them.

Small panels pass gates

The same rule holds at the small end. A smoke run of three cases scored 1.000, with a 95% interval of 0.292 to 1.000. A gate that compares the point estimate with a target ships it. A gate that reads the interval says what is true: three cases cannot tell a perfect system from one that is right three times in ten.

Narrowing the interval

More cases narrow an interval, and so does losing fewer of them. On the same panel, declaring the system's HTTP 500s transient, so that the runner retried them, recovered 18 of 22 missing cases and narrowed the interval by 39%, minutes later, with nothing else changed.

A delta whose interval clears the threshold supports a decision, whatever its size, and that is the one a gate should act on. Until then, report the delta, its interval and the cases behind it, and say the answer is not in yet.

Want this computed for you on every change?

Contact us