Guides
What works today
What you can evaluate with Oloproof today, how you reach it (Python SDK, oloproof.yaml and the CLI, or the browser), and what is not built. It is drawn from the capability audit of 8 October 2026, an internal review that checked each line against the code and its tests rather than against plans.
How to read this page
"SDK" means the Python package (import oloproof). "YAML" means a project you run with oloproof run from oloproof.yaml. They are not the same surface: a few evaluators exist in only one of them, and this page says which. "Browser" means the hosted workbench, which shows what you oloproof push; it does not run evaluations you have not defined in code or YAML.
Oloproof calls your system; it does not host, sandbox or reset it. Your application owns its own state, sessions and tool side effects during a run.
By kind of system
| Your system | What works | Where | Not built |
|---|---|---|---|
| Classifier or structured output | Exact match, contains, regex, JSON schema, rubric judge, probability judge, model classifier; pass rates with intervals | SDK and YAML; model_classifier, probability_judge and cascade in YAML only | n/a |
| Text generation, summaries, extraction | The same checks over free text, rubric judges, judge agreement against human labels, pairwise preference | SDK and YAML; preference is the oloproof prefer command | BLEU, ROUGE or embedding similarity (write one with @evaluator) |
| RAG you run as a black box | Hit rate, recall, MRR, nDCG, citation validity, groundedness and citation-support judges, from retrieval artifacts your system records or returns | SDK and YAML | Diagnosis: diagnose needs a staged system |
| RAG built in stages | All of the above, per-stage caching, and oloproof diagnose with gold context, top-k or reranker changes beside a control | SDK @rag_system, YAML system.rag, CLI diagnose | Interventions beyond those three |
| Tool-using agent | Step limits, required and forbidden tools, tool order, loops, constraints, from a recorded trajectory | SDK and YAML | An LLM-judged trajectory quality check |
| Multi-agent system | Routing, tool permissions per agent, handoff limits | SDK and YAML | n/a |
| Agent replay | Re-running a case from a checkpoint to label a step as needed or not | SDK only (replay_case), for a system that supports checkpoints | A CLI command |
| Binary classifier model | Accuracy, precision, recall, Brier, log loss, ranking (AUC) | SDK and YAML predictive: | n/a |
| Multiclass model | One block per class (one against the rest) | SDK and YAML | Macro or micro averages |
| Regression model | Absolute error within a declared target_range | SDK and YAML | Squared error, R squared, unbounded error |
| Multi-turn conversation | Whether every scripted turn was answered, and a rubric judge over the whole conversation | SDK only (ConversationCompleted, ConversationJudge) | YAML types, turn-level scores, a user simulator, conversation replay |
| Images, audio, video | Nothing native: a case may carry a URL or encoded file through your system, and @evaluator can check the output | n/a | Judges see JSON text only; no media artifacts or rendering |
A multi-turn workaround: make each turn a case and give the turns of one conversation the same group_id, so the analysis treats them as a cluster (Clusters). Each turn is then a separate call, not one recorded dialogue.
Connecting your system
- Python callable: @system in the SDK or system.callable: module:function in YAML. It receives the case input and returns a dictionary. Sync and async functions both work.
- HTTP: system.http with url, method, output_path, artifacts and timeout_s. The case input is sent as the JSON body and the response read as JSON. Headers and authentication cannot be configured yet; put an endpoint that needs a key behind a callable that adds it.
- Custom evaluators: @evaluator in the SDK. oloproof.yaml cannot name one yet.
Workflows
| Workflow | What works | Not built |
|---|---|---|
| Compare two versions | Paired superiority, non-inferiority and equivalence; slices with multiplicity control | n/a |
| CI | oloproof gate exit codes from your policy's block_on, sign-off, signed records; a pull request summary (--summary markdown) | n/a |
| Human review | Labels from a file or the terminal; a browser review queue with assigned reviewers on a hosted workspace | A browser review queue on a local project |
| Browser workbench | Runs, cases, comparisons, diagnostics, evaluators, review and traffic, after oloproof push | Dataset import, judge authoring, schedules, alerts and postmortems (shown as planned) |
| Local browser | n/a | No CLI command serves the workbench locally; it is hosted |
What has been tested, and how
Every family above has automated tests, run on fixtures and mocked providers. Runs against real systems with real models are recorded for RAG, a single agent and a multi-turn conversation; none yet for predictive models or multi-agent systems. Statistical methods that decide releases are each admitted by their own audit; a metric with no admitted interval does not decide a rule on an interval, and a rule that would need one returns MANUAL_REVIEW or INSUFFICIENT_EVIDENCE with its reason.