Changelog
What shipped, and what it changes for a release decision.
Oloproof is on PyPI
oloproof 0.1.0a1 is published, so the golden path starts with pip install oloproof in a clean environment and goes straight to oloproof init and oloproof run. Releases are built and published from a tagged commit in CI, through PyPI trusted publishing, with no stored token.
Replay how a run reached its decision
A run or a comparison that stops early now records every look it took, and the workbench replays them: the anytime-valid interval narrowing at each look against a threshold that does not move, and the look at which every rule decided. It draws only intervals the engine recorded. A fixed-sample interval is never animated as if it were live, since watching one until it looks good voids its guarantee.
A command palette, keyboard navigation and the evidence chain
⌘K or Ctrl+K lists every destination you can see, and a pasted run id or comparison digest opens it. g then a letter navigates, j and k move through a table's rows, and ? lists every shortcut. A case opens beside its run, with its evidence chain and the people who labelled it. Every animation stops under reduced motion.
The hosted workbench on oloproof.com
The workbench now serves oloproof.com, and the end-to-end golden path passed on the host: a reviewer labelled 80 cases the workspace drew, found 7 wrong answers the judge had passed, and the workspace decided Pass on its verified sample (91.2%, interval 74.5% to 100.0%) where the pushed run alone had insufficient evidence. Access is by invitation.
OpenTelemetry traces into oloproof collect
oloproof collect receives OTLP/HTTP from a stock OpenTelemetry SDK, assembles GenAI and OpenInference spans into content-addressed traces, and keeps their content on your machine; a push follows your egress policy. oloproof traces promote turns a trace into a test case with its provenance, and refuses duplicates.
Decisions on samples of production traffic
Your collector signs a commitment to every trace it sealed each hour. The workspace then draws a sample with its own randomness and checks each drawn trace against the commitment, so a sample cannot be curated by withholding traces. oloproof traces evaluate judges the sample on its recorded outputs, the workspace decides it under the evaluation an Owner pinned, and a confidence sequence follows each metric across samples on the project's Production page.
A review queue in the browser
An Owner or Admin opens a queue and assigns reviewers, who label from the keyboard: pass or fail, a score on a declared scale, or a blind preference between two runs, each recorded by account. The reviewer's browser fetches case content from oloproof collect on your side on a five-minute signed ticket, so the workspace never holds it.
Verified execution
A run can be signed by a runner key an Owner registered, or by the managed worker. An Owner pins the whole evaluation (suite, evaluators, metrics and baseline run) and its gate policy, and the workspace decides each system version on its first run of exactly that evaluation, against the pinned run. Human labels count only from independent labellers, counted by account.
Human samples the workspace draws
A PPI interval corrects a judge with a random sample of human labels, and its guarantee needs a sample nobody chose. The workspace now draws that sample after the run's evidence is frozen there, keeps the seed to itself, and checks the corrected interval with its own engine. The run page says whether the workspace verified an interval; a sample drawn on your own machine stays labelled good faith.
Early stopping, and judge rates corrected by people
A policy with early_stopping: true runs cases in seeded batches and stops a run, or a candidate and its baseline in lockstep, once every rule has decided, recorded as DECIDED_EARLY with the cases it did not need. Stopped rates use an anytime-valid betting interval, so looking after every batch keeps its guarantee. A judge's pass rate can be gated on a PPI interval that combines the judge with a blind random sample of human labels, shown beside the judge-only and human-only intervals.
Notifications by email
Four notifications, each a category a member can turn off from the email itself: a managed job finished or failed, a pushed run or comparison whose gate blocks a release, usage at 80% and 100% of the free allowance, and a member joined. A gate email quotes the stored decision and links to the run, with no case content. Email only: no Slack, pager or webhook.
Passwords, and deleting your account
Sign in with an email address and a password as well as with a provider, with verified addresses and resets. A person can delete their own account: the user and every workspace only they belonged to go at once, and the worker purges those workspaces' evidence.
Probability judges, calibration and a cascade
A judge can answer a typed question (yes or no, a choice, a score against levels) with a probability for every answer, read from a local model's token probabilities in one forward pass. Its calibration is measured against human labels and recorded on its registry entry, and a cascade sends only the cases where a cheap judge is unsure to a stronger one. A trained classifier can be an evaluator too, with no key and no network.
Judges held to human labels
Label a run's cases from a file or the terminal, and try a draft judge against those labels before adopting it. A judge reports its bias beside its agreement, one whose bias exceeds a declared margin is barred from gating, and a validation stops counting once the model it measured changes. Label-free probes check whether a verdict moves when nothing that matters did, and pairwise comparisons are asked in both orders, so an answer preferred only for its slot shows up without anyone labelling.
Multi-agent evidence
A trajectory records which agent took each step, and every hand-off. Evaluators judge routing (did the request reach the agent that should handle it), tool permissions (did any agent call a tool that was not its to call) and agents passing control back and forth instead of finishing, and slices group cases by route.
Replicates and a flaky count
A suite can measure each case several times. Replicates are aggregated per case before any interval, comparisons pair them as fractions, and the run reports how many cases disagreed with themselves. A system can mark a failure as transient so the runner retries it: on a live RAG system that recovered 18 of 22 missing cases and narrowed the interval by 39%.