Skip to content

About

We build the measurement layer.

Teams ship changes to AI systems every week with no way to say whether the last one made them better. Oloproof is the part of the stack that answers that question with evidence: an estimate, its uncertainty, and a decision you can check.

How we decide

On the interval, not the estimate

A rule passes when the whole interval clears its threshold and fails when the whole interval misses it. Anything between is reported as insufficient evidence, never rounded to a pass.

What we claim

Nothing we could not reproduce

A statistical method enters the product with a written note, a coverage grid and an audit. Until then, a rule that would need it fails closed rather than borrow it.

Where evidence lives

With you, by default

Local runs never leave your machine. A push carries metrics, intervals and decisions, and leaves raw inputs, outputs and traces at home unless your project says otherwise.

Judges

Measured before they gate

An LLM judge is checked against human labels, on the lower bound of its agreement, before it may decide a release.

Language models

Above the evidence, not inside it

Models may judge, explain and suggest the next experiment. They never invent a statistical method or make the release decision.

Cost

Never pay twice for the same evidence

Every execution and judgment is stored once, by its content. Running an unchanged suite again reuses what it already has.

What it is for

Does it work? For the scenarios and criteria you declared, with the denominator stated.

Where does it fail? Including the slices and components a global rate hides.

How sure are we? Every number that drives a decision carries its uncertainty.

Can it be defended later? The record behind a result is kept, so someone else can check it.

Get in touch