Skip to content

Guides

Results and execution reference

What a run sends to an HTTP system and expects back, how every case enters a metric's denominator, the decision states and the reason codes that explain them, the exit codes, and what data stays local or moves to a hosted workspace. For the ideas behind them read Concepts; for the policy fields read the Configuration reference.

The HTTP system contract

An HTTP system (system.http in oloproof.yaml) is called once per case, and once per replicate.

AspectBehaviour
RequestPOST by default (GET and PUT are accepted). The body is the case's input value as JSON.
Headers and authenticationNone can be configured. The request carries only the HTTP client's defaults. An endpoint that needs a key belongs behind a Python callable system that adds it.
ResponseMust be JSON. output_path selects the output by a dotted path, such as result.answer; without it the whole body is the output. A missing output_path field is recorded as an execution error for that case.
ArtifactsEach http.artifacts entry reads a dotted path from the response. A declared field missing from a response is a contract error, and the run stops with exit 2.
Timeouthttp.timeout_s per request, 30 seconds by default.
RetriesTimeouts, connection failures and HTTP 408, 429 and 5xx are retried, up to four attempts in all, with jittered exponential backoff that honours Retry-After. Other 4xx responses are not retried.
After the last attemptThe case's execution is recorded as an error and the case counts as missing (or as failed, under on_execution_error: fail). The run continues.
ConcurrencyAt most concurrency.system requests in flight, 8 by default.

The URL, method, output path and artifact mapping enter the system's version, but what the server does does not. Change system.version whenever the server's behaviour changes; see the Configuration reference for why.

Three different vocabularies

A result has three kinds of state, and they never stand in for one another.

KindValuesAnswers
Decision statePASS, FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEWWhat the evidence says about one rule.
Execution stateRun status RUN_ERROR or CANCELLED; run completeness PARTIAL; case execution ERROR or TIMEOUTWhat happened to the run or to one case's call. Not a quality result.
Release actionALLOW, WARN, BLOCKWhat your policy does with each decision: block_on states block, warn_on states warn, others allow.

A run that ends normally is DECIDED, or DECIDED_EARLY when early stopping ended it. A run interrupted by Ctrl-C or task cancellation is CANCELLED; one whose harness raised is RUN_ERROR. Either leaves the run PARTIAL, keeps the cases that finished, and lets the next run reuse their cached records. See Errors.

For a minimum threshold T and an interval [L, U], a rule is PASS when L >= T, FAIL when U < T, and INSUFFICIENT_EVIDENCE otherwise. A maximum threshold is symmetric. Before reading the interval, a rule checks whether it should decide at all: first the reasons for MANUAL_REVIEW, then those for INSUFFICIENT_EVIDENCE. The first tier with a reason decides, and lists every reason it found.

Reason codes

Every decision carries one or more reason codes.

MANUAL_REVIEW

CodeMeaning
policy_requires_reviewThe rule sets requires_manual_review: true.
unsupported_methodNo admitted interval exists for this metric in this situation. See below.
unsupported_dependence_structureThe suite declares clusters (group_id) and no admitted method handles them for this metric.
approximate_method_not_permittedThe only interval is approximate, and the policy does not set allow_approximate_methods: true.
evaluator_retiredAn evaluator behind the metric was retired.

INSUFFICIENT_EVIDENCE

CodeMeaning
no_observationsNo case was observed for this metric.
missingness_exceeds_policyMore of the eligible cases are missing than the rule's max_missing_fraction allows.
missingness_unboundedThe method drops missing cases rather than bounding them, and the rule declares no max_missing_fraction.
evaluator_not_validatedA model judge behind the metric has not been validated against human labels, and require_validated_evaluators is on (the default).
evaluator_recalibration_requiredThe judge was validated on a served model this run's verdicts did not come from.
interval_unavailableThe metric has no interval to read.
insufficient_clustersFewer clusters than the policy's min_clusters.
interval_monte_carlo_uncertainThe threshold falls inside the simulation uncertainty of a clustered bound.
interval_overlaps_thresholdThe interval contains the threshold. More cases would narrow it.
interval_unboundedThe interval has no bound on the side the rule reads.
missing_could_change_outcomeAn observed_count rule: the missing cases could take the failures past max_failures.
interval_overlaps_zero, interval_overlaps_margin, interval_overlaps_marginsA comparison whose difference interval straddles zero or a margin.
insufficient_supportA slice comparison rule whose slice has fewer cases than its min_support.
family_correction_withheldA rule in a families: entry that the Holm correction stopped before it.

PASS and FAIL

CodeState
lower_bound_meets_minimum, upper_bound_meets_maximumPASS
upper_bound_below_minimum, lower_bound_above_maximumFAIL
observed_failures_within_limitPASS
observed_failures_exceed_limitFAIL
difference_above_zero, lower_bound_above_margin, interval_within_marginsPASS (comparison)
difference_below_zero, upper_bound_below_margin, interval_outside_marginsFAIL (comparison)
cost_ceiling_exceededSaid beside the state of a cost rule whose declared ceiling a recorded execution exceeded

Hosted workspace only

A workspace that decides a pushed run itself can withhold a decision with ppi_not_verified (it did not verify the interval the decision rests on), execution_not_verified (the outputs did not come from a registered runner) or workspace_cannot_decide (it holds no copy of the policy, or could not read the evidence). See Gating.

How each case enters the denominator

Every metric reports four counts: n_total (cases in the suite), n_eligible, n_observed and n_missing, with n_eligible = n_observed + n_missing. Cases outside n_eligible are listed under exclusions with a reason.

What happened to the caseCounts asIn the denominator
The evaluator returned pass or failobserved, success or failureyes
The evaluator declared itself not applicable (for example, no expected value to compare)excluded, with the reasonno
The system call errored or timed outmissing, or failure under on_execution_error: failyes
The evaluator raised, or a judge's reply could not be readmissingyes
The case never ran because the run was interruptedmissing, and the run is PARTIALyes

A missing case is bounded, not dropped. For a pass rate the interval's lower bound treats every missing case as a failure and its upper bound as a success, so a run with many missing cases has a wide interval that cannot pass a demanding rule; a bounded mean substitutes the ends of its declared range the same way. A method that cannot bound missing cases (a ranking statistic, for example) drops them and records the assumption, and a rule over it reads missingness_unbounded until it declares max_missing_fraction.

An observed_count rule counts failures over the executed suite and reads no interval. It passes only when the observed failures plus every missing case still fit within max_failures.

Metrics with no admitted interval

A rule decides only on an interval whose method has been admitted by audit. Where none exists, the metric is still computed and shown, and a rule over it does not borrow an unvalidated method:

SituationWhat a rule over it reads
A score (mean) metric with no declared range, such as a custom score evaluator without score_rangeMANUAL_REVIEW, unsupported_method
A mean, quantile, ranking or cost metric on a suite that declares group_idMANUAL_REVIEW, unsupported_dependence_structure
A pass rate on a clustered suitean approximate interval: MANUAL_REVIEW unless allow_approximate_methods: true, then the cluster checks above
A quantile or ranking metric with replicates above 1MANUAL_REVIEW, unsupported_method
Any metric on a suite with both group_id and replicatesMANUAL_REVIEW, unsupported_dependence_structure
A comparison on a clustered suiteMANUAL_REVIEW
A slice below min_slice_supportno interval, but slices never reach the gate
human_score, human_preference or cost_per_accepted metricsrefused when the file is read, exit 2

Exit codes

oloproof gate, oloproof run with a policy, and the other commands that decide all use the same codes.

CodeMeaning
0Nothing the policy blocks on: every rule passed, or those that did not are outside block_on.
1A rule in block_on failed.
2The configuration or invocation was wrong, or a system broke its contract; nothing was decided.
3A rule in block_on read INSUFFICIENT_EVIDENCE.
4A rule in block_on read MANUAL_REVIEW.
5The run did not complete and block_on_partial_run is on (the default).

When several apply, the code reported is the first of 1, 5, 4, 3. A state left out of block_on cannot change the exit code: with block_on: [FAIL] and warn_on: [INSUFFICIENT_EVIDENCE], an undecided rule warns and the gate exits 0. Exit 0 therefore means only that nothing your policy blocks on occurred, not that every rule passed. See Gating.

Where the work runs and where the data goes

Local, the default

oloproof run, oloproof gate and the SDK run on your machine. Every record (cases, outputs, artifacts, judgments, metrics and decisions) is written to .oloproof/store.sqlite beside oloproof.yaml, or under OLOPROOF_HOME when set. Nothing is sent to Oloproof. The only network traffic is what your configuration causes: calls to your HTTP system's URL, and calls a model judge or model classifier makes to its provider, which receive the case content they judge and bill you for it.

Pushing to a hosted workspace

oloproof push sends a run's evidence to the workspace you connected with oloproof login. By default it sends metrics, intervals, decisions and aggregate slices, and every record's identity, status, timings and usage, but not its content. Raw content is redacted field by field before anything leaves the machine, and a redacted record says which categories were withheld. A category goes only when egress: in oloproof.yaml lists it:

CategoryWhat it covers
raw_inputsScenario inputs, expected values and case metadata: the dataset rows
raw_outputsWhat the system under test returned for each case
judge_rationalesThe text a judge wrote explaining a verdict, which quotes the output
artifactsRetrieval context, citations and trajectories recorded during a run
error_detailException messages and details, which often carry the input verbatim
system_configThe declared configuration of the system under test and of its evaluators
label_notesThe note a person wrote beside a label, which often quotes the output
span_namesTrace, span, tool and agent names an instrumentation recorded

Record digests are not recomputed after redaction, so a hosted record still names the original evidence, which stays on your machine. Redaction is not encryption, and a metric over a very small slice can still identify the cases behind it.

When reviewers label cases in the hosted review queue, their browser fetches case content from oloproof collect running on your side; it does not pass through the workspace. Provider keys a workspace uses are stored with oloproof credentials set, and oloproof credentials list shows their names, never their values. A managed job runs on a worker Oloproof operates, which does not run your Python code; oloproof job reports its result and exits on its gate.

In a hosted workspace the engine never reuses cached executions, judgments or analyses, because a push can write those caches; it recomputes them.