Guides
SDK reference
Every name the oloproof and oloproof.evaluators packages export, with its signature, whether it is synchronous or asynchronous, and what it returns. For a guided introduction read The Python API first.
Only these two packages are the public surface. Anything imported from oloproof_core is engine internals and can change without notice. Every function below runs locally against the project's store; none of them sends data anywhere unless an evaluator you pass calls a model provider.
Running an evaluation
evaluate and aevaluate
def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult| Argument | Type | What it is |
|---|---|---|
| system | a @system function, a @rag_system class or instance, or a callable | The system under test. |
| dataset | path | A JSONL suite. See Suites. |
| evaluators | list | Instances from oloproof.evaluators, or @evaluator functions. |
| policy | path to a release.yaml, a ReleasePolicy, or None | The release policy. None runs no gate: result.gate is None and nothing is decided. |
| concurrency | ConcurrencyConfig or a mapping such as {"system": 8, "judge": 4} | Calls in flight at once. |
| slices | list of strings | Exploratory slices, as in oloproof.yaml. |
| min_slice_support | integer | Below this many eligible cases a slice has no interval. Default 30. |
| replicates | integer | Measure each case this many times. Default 1. |
evaluate is synchronous. Called with no event loop running it uses asyncio.run; called from inside a running loop (a notebook, an async test) it runs the evaluation on a separate thread and blocks until it finishes, so it is safe in both places. aevaluate is the coroutine; await it from async code.
The remaining keyword arguments (metrics, store, predictive, event_sink, retry_policy, traffic_draw_id) take engine types from oloproof_core and are not part of the stable surface.
from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator
@system(name="support-bot", version="1")
def answer(case):
current_case().usage(input_tokens=12, output_tokens=3)
return {"label": "refund" if "refund" in case["question"].lower() else "other"}
@evaluator(criterion="short_label")
def short_label(case):
return len(case.output["label"]) <= 6
result = evaluate(
system=answer,
dataset="cases.jsonl",
evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)Run on a two-case suite with no policy, this printed:
correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0EvaluationResult
| Member | Type | What it is |
|---|---|---|
| run | run record | The stored run, with its id, status and completeness. |
| suite, system, evaluators | version records | The exact versions this run measured. |
| metrics | tuple of metric results | One per criterion and declared metric: metric, estimate, interval, n_total, n_eligible, n_observed, n_missing, exclusions, method. |
| gate | gate result or None | With a policy: release_action, exit_code, decisions and reasons. |
| decisions | tuple | The gate's decisions, or empty without a policy. |
| cases() | list | Every case with its execution and judgments. |
| failures() | list | Cases that did not finish, errored, or failed at least one evaluator. |
| print(stderr=False) | none | The terminal report oloproof run prints. |
| to_bundle(path) | path | Writes a portable bundle, as oloproof export does. |
What the states, reason codes and counts mean is in Results and execution.
evaluate_comparison and aevaluate_comparison
def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResultRuns both systems on the same suite in one seeded order and decides the policy's comparison rules. policy is required and must contain at least one superiority, non_inferiority or equivalence rule, or the call raises a configuration error. Extra keyword arguments: concurrency, replicates, and the engine-typed metrics, store and retry_policy. Synchronous and asynchronous behaviour are as for evaluate.
ComparisonEvaluationResult has candidate and baseline (each an EvaluationResult) and comparison, which carries the paired differences and their decisions. See Comparing two versions.
Declaring a system
@system
def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())Usable bare (@system) or with arguments (@system(name=..., version=...)), or called on an object (system(model.answer, version="v2")). The function receives the case's input object, not the whole case, and returns the output evaluators read. It may be def or async def; a synchronous function runs on a worker thread.
| Argument | Default | What it is |
|---|---|---|
| name | the function's name | Part of the version identity. |
| version | absent | Required for a bound method or a callable object, because their behaviour depends on state Oloproof cannot see. |
| config | empty | Settings recorded with the version. |
| timeout_s | 120 | Per-call limit. A call that exceeds it is recorded as a timed-out execution. |
| records | empty | Artifact kinds the system records, such as retrieval/v1. An evaluator that requires a kind the system does not declare is refused before the run starts. |
The function's own module source enters the version digest, so editing it invalidates cached executions. See Configuration reference for what else does and does not.
current_case
def current_case() -> CaseRecorderAvailable only while Oloproof is calling your system; anywhere else it raises RuntimeError. The recorder's methods:
| Method | Records |
|---|---|
| usage(*, input_tokens=None, output_tokens=None, cost_usd=None) | Tokens and cost of a model call. A value left out stays unrecorded, not zero. |
| artifact(kind, data) | Any JSON value or Pydantic model under a kind such as trace or conversation/v1. |
| retrieval(retrieval) | The ranked candidates a retriever returned (retrieval/v1). |
| context(context) | The context assembled for generation (context/v1). |
| citations(ids) | The ids an answer cites, as doc_id or doc_id#chunk_id (citations/v1). |
| agent_trajectory(trajectory) | An agent's steps, tool calls and results, and checkpoints (agent_trajectory/v1). |
Each returns an ArtifactRef (except usage, which returns nothing). The typed payloads are exported for building these records: Retrieval, Passage, Context, ContextItem, DroppedItem, Citations, StageTimings, AgentTrajectory, AgentStep, AgentCheckpoint, AgentConstraintCheck, and the kind name CONVERSATION (conversation/v1).
@rag_system
def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")A class decorator. The class provides retrieve(input, depth) and generate(input, context), plus count_tokens(passage) when it sets a token_budget. context is the list of Passage objects that survived top_k and the budget, in rank order. Oloproof records retrieval/v1, context/v1, citations/v1 and stage_timings/v1 itself, and caches each stage separately. citations_path names the output field holding the ids the answer cites. See RAG.
Evaluators
All classes are in oloproof.evaluators. Each one's criterion names the metric it produces. Which artifacts each reads, and its YAML equivalent, are in the evaluator table of the Configuration reference.
| Class | Signature |
|---|---|
| ExactMatch | (*, criterion, field=None, expected_field=None, strip=True, casefold=False) |
| Contains | (*, criterion, field=None, expected_field=None) |
| Regex | (*, criterion, pattern, field=None, pass_if="match") |
| JsonSchema | (*, criterion, schema, field=None) |
| RubricJudge | (*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0) |
| Groundedness | (*, provider, model, criterion="groundedness", **options) |
| CitationSupport | (*, provider, model, criterion="citation_support", **options) |
| CitationValidity | (*, criterion="citations_valid", require_citations=False) |
| HitRate, Recall | (k=None, *, criterion=None, relevance_unit="doc"), k defaults to 5 |
| MRR, NDCG | (k=None, *, criterion=None, relevance_unit="doc"), k defaults to 10 |
| AgentMaxSteps | (max_steps, *, criterion=None) |
| AgentToolCalled | (tool_name, *, min_calls=1, criterion=None) |
| AgentNoToolLoop | (*, max_repeats=2, criterion="agent_no_tool_loop") |
| AgentToolSequence | (*, ordered=True, criterion="agent_tool_sequence") |
| AgentNoUndeclaredTool | (*, criterion="agent_no_undeclared_tool") |
| AgentConstraintsSatisfied | (constraints=(), *, criterion="agent_constraints_satisfied") |
| AgentRoute | (*, criterion="agent_route") |
| AgentToolPermissions | (permissions, *, criterion="agent_tool_permissions") |
| AgentMaxHandoffs | (max_handoffs, *, criterion=None) |
| ConversationCompleted | (*, criterion="conversation_completed") |
| ConversationJudge | as RubricJudge |
| PredictiveCorrect, PredictiveRecall, PredictivePrecision | (*, criterion, positive=True, field="label", expected_field="label") |
| AbsoluteError | (*, criterion, target_range, field="label", expected_field="label") |
| Brier, PredictiveRanking | (*, criterion, positive=True, field="score", expected_field="label") |
| LogLoss | (*, clip, criterion, positive=True, field="score", expected_field="label") |
| CustomEvaluator | (func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None) |
provider is "anthropic", "openai" or "openai_compatible". A judge reads its key from the environment variable api_key_env names (ANTHROPIC_API_KEY or OPENAI_API_KEY by default) and is billed by that provider. Groundedness and CitationSupport take the rest of RubricJudge's settings through **options. The probability judge, the model classifier and the cascade have no SDK class; they exist in YAML only.
ConversationCompleted and ConversationJudge read a conversation/v1 artifact your system records. Oloproof does not drive the conversation: your application runs every turn and records the transcript. See Agents.
@evaluator
def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)Wraps a function of one argument, the case, into a CustomEvaluator. The case has output, expected and scenario, and artifacts(name) returns the payloads of a recorded kind. The function may be def or async def. A binary evaluator returns True or False; a score evaluator declares value_type="score" and score_range=(low, high) and returns a number. An exception raised by the function records the case as missing for that criterion, never as a failure.
reads must list every field the function reads (input, output, expected, metadata, metadata.<key> or artifacts.<name>), because the cached judgment is keyed on exactly those. Judgments are reused across runs only with cacheable=True. The defining module's source enters the version, so editing it invalidates them. YAML cannot name a custom evaluator.
Diagnosis
diagnose and adiagnose
def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResultRe-executes a stored run's failed cases under one intervention: "gold-context", "top-k" (with top_k) or "reranker" (with reranker). system and evaluators must be the versions the run used; a different version is refused before anything executes. With control=True a fresh control sample runs beside the intervention, so a change can be told apart from run-to-run variation. diagnose is the synchronous form and behaves inside a running loop as evaluate does. See RAG for the workflow.
InterventionResult holds the parent run id, the intervention, whether it was supported, the intervention and control runs, a per-case outcome, and a DiagnosisReport of CaseDiagnosis entries when one was produced.
Agent replay
def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]Replay is something your system does, not something Oloproof simulates. A system supports it only by implementing async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory, returning what the agent did from the checkpoint onward; Oloproof splices on the recorded prefix and compares. Your application owns its state, sessions and tool side effects, including resetting them before a replay. supports_replay reports whether a system declares the method.
replay_case is a coroutine: await it, or call it through asyncio.run. It runs a control (the same checkpoint with ReplayChange(kind="resume")) before the replay, and does not attempt the replay when the control does not reproduce the recording. change is ReplayChange(kind="drop_step", step_index=...) or ReplayChange(kind="resume"). checkpoint_for picks the latest recorded checkpoint strictly before a step, or None, in which case the outcome is discarded as no_checkpoint_recorded without touching the system.
ReplayOutcome carries scenario_id, change, the terminal statuses of the replay and control, and discarded with a reason when the case yields no evidence. A discarded case is not a failed case. label_case turns an outcome into a CaseDiagnosis with a FailureLabel and its LabelReason; unnecessary_steps lists the steps whose removal left the outcome intact. See Agents.
Human labels and evaluator trust
def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabelStores one person's pass/fail verdict on one case of a stored run. Labels feed oloproof evaluators validate, which measures a judge's agreement against them (AgreementResult) and records its status (RegistryEntry). The further measurement_sample_* arguments bind a label to a measurement sample; the CLI's oloproof labels export and oloproof labels import fill them for you. See Judges.
Policies in code
ReleasePolicy, IntervalThresholdRule, ObservedCountRule and DecisionRule (the union of the two) build a policy without a file. Their fields are the release.yaml fields of the Configuration reference; an interval rule takes direction (min or max) and threshold instead of min: or max:. Comparison rules have no exported class; write them in release.yaml and pass its path.
from oloproof import IntervalThresholdRule, ReleasePolicy
policy = ReleasePolicy(
rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)Other exports
| Name | What it is |
|---|---|
| ConcurrencyConfig | system and judge limits, as in oloproof.yaml. |
| TransientError | Raise it from a system, with retryable=True, to have the call retried with backoff. |
| ArtifactRef | The reference a recorded artifact returns: its kind and digest. |
| __version__ | The installed package version. |