Skip to content

Guides

Start here

What Oloproof does, when to use it and when not to, and the order to read these guides in: from a first evaluation of one application to comparing versions, review, CI and the exact contracts.

What Oloproof does

Oloproof runs your AI application over a fixed set of cases, checks each output with evaluators you choose, and turns the results into rates with intervals. A release policy you write then decides each rule as PASS, FAIL, INSUFFICIENT_EVIDENCE or MANUAL_REVIEW, and the command exits on that decision so CI can use it. Everything runs on your machine; nothing leaves it unless you push a run to a hosted workspace. The terms are defined in Core concepts.

Use it when you need to know whether an application meets a stated bar on cases you trust, how sure that answer is, and whether a change made it better or worse. Oloproof calls your system; it does not host, sandbox or reset it, and your application owns its own state, sessions and tool side effects. For what is and is not built, by kind of system, read What works today before you commit to it.

These guides describe Oloproof 0.1.0a5. pip show oloproof prints the version you have installed; an older one may lack a command or flag shown here, such as oloproof init --example, and pip install --upgrade --pre oloproof brings it up to date.

The journey

StepWhat you doRead
1Learn what Oloproof does and when to use itThis page, Core concepts, What works today
2Evaluate your first application, without a baselineQuickstart
3Choose a guide for your system typeThe table below
4Inspect failures and understand the resultSuites, Errors, Clustered cases, Slices
5Change the system and compare versionsCompare, Comparison rules
6Add human review, collaboration and CIJudges, Gating, CI, Notifications
7Look up exact technical contractsConfiguration, SDK reference, Results, CLI

A first evaluation needs no baseline: it asks whether one version of your application meets your requirements. Comparing two versions comes at step 5, once you have a run worth comparing against.

Choose your system type

Your systemStart with
A classifier, router or structured output (a label, a JSON object)Classification
Free text: answers, summaries, extractionText generation
A service you reach over HTTP, in any languageAn HTTP endpoint
Retrieval augmented generationRAG
A tool-using agent, or several agentsAgents
A trained predictive model (classification or regression)Predictive models
A multi-turn conversationConversations

Images, audio and video have no native support: a case can carry a URL or an encoded file through your system, and a custom evaluator can check the output, but judges see JSON text only.