Skip to content

Guides

SDK reference

Every name the oloproof and oloproof.evaluators packages export, with its signature, whether it is synchronous or asynchronous, and what it returns. For a guided introduction read The Python API first.

Only these two packages are the public surface. Anything imported from oloproof_core is engine internals and can change without notice. Every function below runs locally against the project's store; none of them sends data anywhere unless an evaluator you pass calls a model provider.

Running an evaluation

evaluate and aevaluate

def evaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
async def aevaluate(*, system, dataset, evaluators, policy=None, **kwargs) -> EvaluationResult
ArgumentTypeWhat it is
systema @system function, a @rag_system class or instance, or a callableThe system under test.
datasetpathA JSONL suite. See Suites.
evaluatorslistInstances from oloproof.evaluators, or @evaluator functions.
policypath to a release.yaml, a ReleasePolicy, or NoneThe release policy. None runs no gate: result.gate is None and nothing is decided.
concurrencyConcurrencyConfig or a mapping such as {"system": 8, "judge": 4}Calls in flight at once.
sliceslist of stringsExploratory slices, as in oloproof.yaml.
min_slice_supportintegerBelow this many eligible cases a slice has no interval. Default 30.
replicatesintegerMeasure each case this many times. Default 1.

evaluate is synchronous. Called with no event loop running it uses asyncio.run; called from inside a running loop (a notebook, an async test) it runs the evaluation on a separate thread and blocks until it finishes, so it is safe in both places. aevaluate is the coroutine; await it from async code.

The remaining keyword arguments (metrics, store, predictive, event_sink, retry_policy, traffic_draw_id) take engine types from oloproof_core and are not part of the stable surface.

from oloproof import evaluate, system, current_case
from oloproof.evaluators import ExactMatch, evaluator

@system(name="support-bot", version="1")
def answer(case):
    current_case().usage(input_tokens=12, output_tokens=3)
    return {"label": "refund" if "refund" in case["question"].lower() else "other"}

@evaluator(criterion="short_label")
def short_label(case):
    return len(case.output["label"]) <= 6

result = evaluate(
    system=answer,
    dataset="cases.jsonl",
    evaluators=[ExactMatch(criterion="correct_label", field="label"), short_label],
)
for metric in result.metrics:
    print(metric.metric, metric.estimate, metric.interval, metric.n_observed, metric.n_missing)

Run on a two-case suite with no policy, this printed:

correct_label 1.0 lower=0.15811388300841903 upper=1.0 2 0
short_label 1.0 lower=0.15811388300841903 upper=1.0 2 0

EvaluationResult

MemberTypeWhat it is
runrun recordThe stored run, with its id, status and completeness.
suite, system, evaluatorsversion recordsThe exact versions this run measured.
metricstuple of metric resultsOne per criterion and declared metric: metric, estimate, interval, n_total, n_eligible, n_observed, n_missing, exclusions, method.
gategate result or NoneWith a policy: release_action, exit_code, decisions and reasons.
decisionstupleThe gate's decisions, or empty without a policy.
cases()listEvery case with its execution and judgments.
failures()listCases that did not finish, errored, or failed at least one evaluator.
print(stderr=False)noneThe terminal report oloproof run prints.
to_bundle(path)pathWrites a portable bundle, as oloproof export does.

What the states, reason codes and counts mean is in Results and execution.

evaluate_comparison and aevaluate_comparison

def evaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult
async def aevaluate_comparison(*, candidate_system, baseline_system, dataset, evaluators, policy, **kwargs) -> ComparisonEvaluationResult

Runs both systems on the same suite in one seeded order and decides the policy's comparison rules. policy is required and must contain at least one superiority, non_inferiority or equivalence rule, or the call raises a configuration error. Extra keyword arguments: concurrency, replicates, and the engine-typed metrics, store and retry_policy. Synchronous and asynchronous behaviour are as for evaluate.

ComparisonEvaluationResult has candidate and baseline (each an EvaluationResult) and comparison, which carries the paired differences and their decisions. See Comparing two versions.

Declaring a system

@system

def system(func=None, *, name=None, version=None, config=None, timeout_s=None, records=())

Usable bare (@system) or with arguments (@system(name=..., version=...)), or called on an object (system(model.answer, version="v2")). The function receives the case's input object, not the whole case, and returns the output evaluators read. It may be def or async def; a synchronous function runs on a worker thread.

ArgumentDefaultWhat it is
namethe function's namePart of the version identity.
versionabsentRequired for a bound method or a callable object, because their behaviour depends on state Oloproof cannot see.
configemptySettings recorded with the version.
timeout_s120Per-call limit. A call that exceeds it is recorded as a timed-out execution.
recordsemptyArtifact kinds the system records, such as retrieval/v1. An evaluator that requires a kind the system does not declare is refused before the run starts.

The function's own module source enters the version digest, so editing it invalidates cached executions. See Configuration reference for what else does and does not.

current_case

def current_case() -> CaseRecorder

Available only while Oloproof is calling your system; anywhere else it raises RuntimeError. The recorder's methods:

MethodRecords
usage(*, input_tokens=None, output_tokens=None, cost_usd=None)Tokens and cost of a model call. A value left out stays unrecorded, not zero.
artifact(kind, data)Any JSON value or Pydantic model under a kind such as trace or conversation/v1.
retrieval(retrieval)The ranked candidates a retriever returned (retrieval/v1).
context(context)The context assembled for generation (context/v1).
citations(ids)The ids an answer cites, as doc_id or doc_id#chunk_id (citations/v1).
agent_trajectory(trajectory)An agent's steps, tool calls and results, and checkpoints (agent_trajectory/v1).

Each returns an ArtifactRef (except usage, which returns nothing). The typed payloads are exported for building these records: Retrieval, Passage, Context, ContextItem, DroppedItem, Citations, StageTimings, AgentTrajectory, AgentStep, AgentCheckpoint, AgentConstraintCheck, and the kind name CONVERSATION (conversation/v1).

@rag_system

def rag_system(*, name, depth, top_k, token_budget=None, index_version=None, version=None, config=None, citations_path="citations")

A class decorator. The class provides retrieve(input, depth) and generate(input, context), plus count_tokens(passage) when it sets a token_budget. context is the list of Passage objects that survived top_k and the budget, in rank order. Oloproof records retrieval/v1, context/v1, citations/v1 and stage_timings/v1 itself, and caches each stage separately. citations_path names the output field holding the ids the answer cites. See RAG.

Evaluators

All classes are in oloproof.evaluators. Each one's criterion names the metric it produces. Which artifacts each reads, and its YAML equivalent, are in the evaluator table of the Configuration reference.

ClassSignature
ExactMatch(*, criterion, field=None, expected_field=None, strip=True, casefold=False)
Contains(*, criterion, field=None, expected_field=None)
Regex(*, criterion, pattern, field=None, pass_if="match")
JsonSchema(*, criterion, schema, field=None)
RubricJudge(*, criterion, provider, model, rubric_text=None, rubric_file=None, api_key_env=None, base_url=None, temperature=0, max_tokens=512, timeout_s=60.0)
Groundedness(*, provider, model, criterion="groundedness", **options)
CitationSupport(*, provider, model, criterion="citation_support", **options)
CitationValidity(*, criterion="citations_valid", require_citations=False)
HitRate, Recall(k=None, *, criterion=None, relevance_unit="doc"), k defaults to 5
MRR, NDCG(k=None, *, criterion=None, relevance_unit="doc"), k defaults to 10
AgentMaxSteps(max_steps, *, criterion=None)
AgentToolCalled(tool_name, *, min_calls=1, criterion=None)
AgentNoToolLoop(*, max_repeats=2, criterion="agent_no_tool_loop")
AgentToolSequence(*, ordered=True, criterion="agent_tool_sequence")
AgentNoUndeclaredTool(*, criterion="agent_no_undeclared_tool")
AgentConstraintsSatisfied(constraints=(), *, criterion="agent_constraints_satisfied")
AgentRoute(*, criterion="agent_route")
AgentToolPermissions(permissions, *, criterion="agent_tool_permissions")
AgentMaxHandoffs(max_handoffs, *, criterion=None)
ConversationCompleted(*, criterion="conversation_completed")
ConversationJudgeas RubricJudge
PredictiveCorrect, PredictiveRecall, PredictivePrecision(*, criterion, positive=True, field="label", expected_field="label")
AbsoluteError(*, criterion, target_range, field="label", expected_field="label")
Brier, PredictiveRanking(*, criterion, positive=True, field="score", expected_field="label")
LogLoss(*, clip, criterion, positive=True, field="score", expected_field="label")
CustomEvaluator(func, *, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

provider is "anthropic", "openai" or "openai_compatible". A judge reads its key from the environment variable api_key_env names (ANTHROPIC_API_KEY or OPENAI_API_KEY by default) and is billed by that provider. Groundedness and CitationSupport take the rest of RubricJudge's settings through **options. The probability judge, the model classifier and the cascade have no SDK class; they exist in YAML only.

ConversationCompleted and ConversationJudge read a conversation/v1 artifact your system records. Oloproof does not drive the conversation: your application runs every turn and records the transcript. See Agents.

@evaluator

def evaluator(*, criterion, reads=("output", "expected"), cacheable=False, version=None, value_type="binary", score_range=None)

Wraps a function of one argument, the case, into a CustomEvaluator. The case has output, expected and scenario, and artifacts(name) returns the payloads of a recorded kind. The function may be def or async def. A binary evaluator returns True or False; a score evaluator declares value_type="score" and score_range=(low, high) and returns a number. An exception raised by the function records the case as missing for that criterion, never as a failure.

reads must list every field the function reads (input, output, expected, metadata, metadata.<key> or artifacts.<name>), because the cached judgment is keyed on exactly those. Judgments are reused across runs only with cacheable=True. The defining module's source enters the version, so editing it invalidates them. YAML cannot name a custom evaluator.

Diagnosis

diagnose and adiagnose

def diagnose(run_id, **kwargs) -> InterventionResult
async def adiagnose(run_id, *, system, evaluators, intervention, criterion, control=True, top_k=None, reranker=None, reranker_root=None, store=None, concurrency=None) -> InterventionResult

Re-executes a stored run's failed cases under one intervention: "gold-context", "top-k" (with top_k) or "reranker" (with reranker). system and evaluators must be the versions the run used; a different version is refused before anything executes. With control=True a fresh control sample runs beside the intervention, so a change can be told apart from run-to-run variation. diagnose is the synchronous form and behaves inside a running loop as evaluate does. See RAG for the workflow.

InterventionResult holds the parent run id, the intervention, whether it was supported, the intervention and control runs, a per-case outcome, and a DiagnosisReport of CaseDiagnosis entries when one was produced.

Agent replay

def supports_replay(system) -> bool
def checkpoint_for(trajectory, step_index) -> AgentCheckpoint | None
async def replay_case(system, *, scenario_id, trajectory, checkpoint, change) -> ReplayOutcome
def label_case(outcome) -> CaseDiagnosis
def label_cases(outcomes) -> tuple[CaseDiagnosis, ...]
def unnecessary_steps(labels) -> tuple[UnnecessaryStep, ...]

Replay is something your system does, not something Oloproof simulates. A system supports it only by implementing async def replay(self, trajectory, *, checkpoint, change) -> AgentTrajectory, returning what the agent did from the checkpoint onward; Oloproof splices on the recorded prefix and compares. Your application owns its state, sessions and tool side effects, including resetting them before a replay. supports_replay reports whether a system declares the method.

replay_case is a coroutine: await it, or call it through asyncio.run. It runs a control (the same checkpoint with ReplayChange(kind="resume")) before the replay, and does not attempt the replay when the control does not reproduce the recording. change is ReplayChange(kind="drop_step", step_index=...) or ReplayChange(kind="resume"). checkpoint_for picks the latest recorded checkpoint strictly before a step, or None, in which case the outcome is discarded as no_checkpoint_recorded without touching the system.

ReplayOutcome carries scenario_id, change, the terminal statuses of the replay and control, and discarded with a reason when the case yields no evidence. A discarded case is not a failed case. label_case turns an outcome into a CaseDiagnosis with a FailureLabel and its LabelReason; unnecessary_steps lists the steps whose removal left the outcome intact. See Agents.

Human labels and evaluator trust

def record_label(*, run_id, scenario_id, criterion, passed, labelled_by, note=None, purpose="measurement", sample_index=0, config="oloproof.yaml", store=None, ...) -> HumanLabel

Stores one person's pass/fail verdict on one case of a stored run. Labels feed oloproof evaluators validate, which measures a judge's agreement against them (AgreementResult) and records its status (RegistryEntry). The further measurement_sample_* arguments bind a label to a measurement sample; the CLI's oloproof labels export and oloproof labels import fill them for you. See Judges.

Policies in code

ReleasePolicy, IntervalThresholdRule, ObservedCountRule and DecisionRule (the union of the two) build a policy without a file. Their fields are the release.yaml fields of the Configuration reference; an interval rule takes direction (min or max) and threshold instead of min: or max:. Comparison rules have no exported class; write them in release.yaml and pass its path.

from oloproof import IntervalThresholdRule, ReleasePolicy

policy = ReleasePolicy(
    rules=(IntervalThresholdRule(id="accuracy", metric="correct_label", direction="min", threshold=0.8),),
)

Other exports

NameWhat it is
ConcurrencyConfigsystem and judge limits, as in oloproof.yaml.
TransientErrorRaise it from a system, with retryable=True, to have the call retried with backoff.
ArtifactRefThe reference a recorded artifact returns: its kind and digest.
__version__The installed package version.