Guides
Start here
What Oloproof does, when to use it and when not to, and the order to read these guides in: from a first evaluation of one application to comparing versions, review, CI and the exact contracts.
What Oloproof does
Oloproof runs your AI application over a fixed set of cases, checks each output with evaluators you choose, and turns the results into rates with intervals. A release policy you write then decides each rule as PASS, FAIL, INSUFFICIENT_EVIDENCE or MANUAL_REVIEW, and the command exits on that decision so CI can use it. Everything runs on your machine; nothing leaves it unless you push a run to a hosted workspace. The terms are defined in Core concepts.
Use it when you need to know whether an application meets a stated bar on cases you trust, how sure that answer is, and whether a change made it better or worse. Oloproof calls your system; it does not host, sandbox or reset it, and your application owns its own state, sessions and tool side effects. For what is and is not built, by kind of system, read What works today before you commit to it.
These guides describe Oloproof 0.1.0a5. pip show oloproof prints the version you have installed; an older one may lack a command or flag shown here, such as oloproof init --example, and pip install --upgrade --pre oloproof brings it up to date.
The journey
| Step | What you do | Read |
|---|---|---|
| 1 | Learn what Oloproof does and when to use it | This page, Core concepts, What works today |
| 2 | Evaluate your first application, without a baseline | Quickstart |
| 3 | Choose a guide for your system type | The table below |
| 4 | Inspect failures and understand the result | Suites, Errors, Clustered cases, Slices |
| 5 | Change the system and compare versions | Compare, Comparison rules |
| 6 | Add human review, collaboration and CI | Judges, Gating, CI, Notifications |
| 7 | Look up exact technical contracts | Configuration, SDK reference, Results, CLI |
A first evaluation needs no baseline: it asks whether one version of your application meets your requirements. Comparing two versions comes at step 5, once you have a run worth comparing against.
Choose your system type
| Your system | Start with |
|---|---|
| A classifier, router or structured output (a label, a JSON object) | Classification |
| Free text: answers, summaries, extraction | Text generation |
| A service you reach over HTTP, in any language | An HTTP endpoint |
| Retrieval augmented generation | RAG |
| A tool-using agent, or several agents | Agents |
| A trained predictive model (classification or regression) | Predictive models |
| A multi-turn conversation | Conversations |
Images, audio and video have no native support: a case can carry a URL or an encoded file through your system, and a custom evaluator can check the output, but judges see JSON text only.