Comparison
Oloproof vs ad hoc prompt testing.
Spot-checking in a playground is fast and genuinely useful while you are exploring. It stops being enough the moment a change ships to users.
| Question | Playground spot-check | Oloproof |
|---|---|---|
| Did quality change? | A feeling from a handful of outputs | A difference with a 95% interval over the paired cases, and a decision made on the interval |
| For whom did it change? | Unknown | Per declared slice: language, intent or any metadata key, the first tool an agent called, its route |
| Can you rerun it next quarter? | No: the prompts and cases are gone | Yes: the suite, system version and evaluator versions are content-addressed, and a sample keeps its seed |
| Does it block a bad release? | Only if someone remembers to look | The gate exits non-zero and the CI job fails: 1 when a rule fails, 3 when the evidence does not decide |
| What do you show risk or a customer? | Screenshots in a thread | An exported bundle: the run, its suite, system, evaluators, cases and sign-offs |
| Time to first signal | Minutes | Minutes locally, then on every change your CI runs it on |
Where spot-checking still wins
Early exploration, debugging a single strange output, and sanity checks while you write a prompt. Keep doing it. An evaluation suite is what you build once the behaviour matters enough to defend.