Skip to content

Guides

What works today

What you can evaluate with Oloproof today, how you reach it (Python SDK, oloproof.yaml and the CLI, or the browser), and what is not built. It is drawn from the capability audit of 8 October 2026, an internal review that checked each line against the code and its tests rather than against plans.

How to read this page

"SDK" means the Python package (import oloproof). "YAML" means a project you run with oloproof run from oloproof.yaml. They are not the same surface: a few evaluators exist in only one of them, and this page says which. "Browser" means the hosted workbench, which shows what you oloproof push; it does not run evaluations you have not defined in code or YAML.

Oloproof calls your system; it does not host, sandbox or reset it. Your application owns its own state, sessions and tool side effects during a run.

By kind of system

Your systemWhat worksWhereNot built
Classifier or structured outputExact match, contains, regex, JSON schema, rubric judge, probability judge, model classifier; pass rates with intervalsSDK and YAML; model_classifier, probability_judge and cascade in YAML onlyn/a
Text generation, summaries, extractionThe same checks over free text, rubric judges, judge agreement against human labels, pairwise preferenceSDK and YAML; preference is the oloproof prefer commandBLEU, ROUGE or embedding similarity (write one with @evaluator)
RAG you run as a black boxHit rate, recall, MRR, nDCG, citation validity, groundedness and citation-support judges, from retrieval artifacts your system records or returnsSDK and YAMLDiagnosis: diagnose needs a staged system
RAG built in stagesAll of the above, per-stage caching, and oloproof diagnose with gold context, top-k or reranker changes beside a controlSDK @rag_system, YAML system.rag, CLI diagnoseInterventions beyond those three
Tool-using agentStep limits, required and forbidden tools, tool order, loops, constraints, from a recorded trajectorySDK and YAMLAn LLM-judged trajectory quality check
Multi-agent systemRouting, tool permissions per agent, handoff limitsSDK and YAMLn/a
Agent replayRe-running a case from a checkpoint to label a step as needed or notSDK only (replay_case), for a system that supports checkpointsA CLI command
Binary classifier modelAccuracy, precision, recall, Brier, log loss, ranking (AUC)SDK and YAML predictive:n/a
Multiclass modelOne block per class (one against the rest)SDK and YAMLMacro or micro averages
Regression modelAbsolute error within a declared target_rangeSDK and YAMLSquared error, R squared, unbounded error
Multi-turn conversationWhether every scripted turn was answered, and a rubric judge over the whole conversationSDK only (ConversationCompleted, ConversationJudge)YAML types, turn-level scores, a user simulator, conversation replay
Images, audio, videoNothing native: a case may carry a URL or encoded file through your system, and @evaluator can check the outputn/aJudges see JSON text only; no media artifacts or rendering

A multi-turn workaround: make each turn a case and give the turns of one conversation the same group_id, so the analysis treats them as a cluster (Clusters). Each turn is then a separate call, not one recorded dialogue.

Connecting your system

  • Python callable: @system in the SDK or system.callable: module:function in YAML. It receives the case input and returns a dictionary. Sync and async functions both work.
  • HTTP: system.http with url, method, output_path, artifacts and timeout_s. The case input is sent as the JSON body and the response read as JSON. Headers and authentication cannot be configured yet; put an endpoint that needs a key behind a callable that adds it.
  • Custom evaluators: @evaluator in the SDK. oloproof.yaml cannot name one yet.

Workflows

WorkflowWhat worksNot built
Compare two versionsPaired superiority, non-inferiority and equivalence; slices with multiplicity controln/a
CIoloproof gate exit codes from your policy's block_on, sign-off, signed records; a pull request summary (--summary markdown)n/a
Human reviewLabels from a file or the terminal; a browser review queue with assigned reviewers on a hosted workspaceA browser review queue on a local project
Browser workbenchRuns, cases, comparisons, diagnostics, evaluators, review and traffic, after oloproof pushDataset import, judge authoring, schedules, alerts and postmortems (shown as planned)
Local browsern/aNo CLI command serves the workbench locally; it is hosted

What has been tested, and how

Every family above has automated tests, run on fixtures and mocked providers. Runs against real systems with real models are recorded for RAG, a single agent and a multi-turn conversation; none yet for predictive models or multi-agent systems. Statistical methods that decide releases are each admitted by their own audit; a metric with no admitted interval does not decide a rule on an interval, and a rule that would need one returns MANUAL_REVIEW or INSUFFICIENT_EVIDENCE with its reason.