Guides
Results and execution reference
What a run sends to an HTTP system and expects back, how every case enters a metric's denominator, the decision states and the reason codes that explain them, the exit codes, and what data stays local or moves to a hosted workspace. For the ideas behind them read Concepts; for the policy fields read the Configuration reference.
The HTTP system contract
An HTTP system (system.http in oloproof.yaml) is called once per case, and once per replicate.
| Aspect | Behaviour |
|---|---|
| Request | POST by default (GET and PUT are accepted). The body is the case's input value as JSON. |
| Headers and authentication | None can be configured. The request carries only the HTTP client's defaults. An endpoint that needs a key belongs behind a Python callable system that adds it. |
| Response | Must be JSON. output_path selects the output by a dotted path, such as result.answer; without it the whole body is the output. A missing output_path field is recorded as an execution error for that case. |
| Artifacts | Each http.artifacts entry reads a dotted path from the response. A declared field missing from a response is a contract error, and the run stops with exit 2. |
| Timeout | http.timeout_s per request, 30 seconds by default. |
| Retries | Timeouts, connection failures and HTTP 408, 429 and 5xx are retried, up to four attempts in all, with jittered exponential backoff that honours Retry-After. Other 4xx responses are not retried. |
| After the last attempt | The case's execution is recorded as an error and the case counts as missing (or as failed, under on_execution_error: fail). The run continues. |
| Concurrency | At most concurrency.system requests in flight, 8 by default. |
The URL, method, output path and artifact mapping enter the system's version, but what the server does does not. Change system.version whenever the server's behaviour changes; see the Configuration reference for why.
Three different vocabularies
A result has three kinds of state, and they never stand in for one another.
| Kind | Values | Answers |
|---|---|---|
| Decision state | PASS, FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW | What the evidence says about one rule. |
| Execution state | Run status RUN_ERROR or CANCELLED; run completeness PARTIAL; case execution ERROR or TIMEOUT | What happened to the run or to one case's call. Not a quality result. |
| Release action | ALLOW, WARN, BLOCK | What your policy does with each decision: block_on states block, warn_on states warn, others allow. |
A run that ends normally is DECIDED, or DECIDED_EARLY when early stopping ended it. A run interrupted by Ctrl-C or task cancellation is CANCELLED; one whose harness raised is RUN_ERROR. Either leaves the run PARTIAL, keeps the cases that finished, and lets the next run reuse their cached records. See Errors.
For a minimum threshold T and an interval [L, U], a rule is PASS when L >= T, FAIL when U < T, and INSUFFICIENT_EVIDENCE otherwise. A maximum threshold is symmetric. Before reading the interval, a rule checks whether it should decide at all: first the reasons for MANUAL_REVIEW, then those for INSUFFICIENT_EVIDENCE. The first tier with a reason decides, and lists every reason it found.
Reason codes
Every decision carries one or more reason codes.
MANUAL_REVIEW
| Code | Meaning |
|---|---|
| policy_requires_review | The rule sets requires_manual_review: true. |
| unsupported_method | No admitted interval exists for this metric in this situation. See below. |
| unsupported_dependence_structure | The suite declares clusters (group_id) and no admitted method handles them for this metric. |
| approximate_method_not_permitted | The only interval is approximate, and the policy does not set allow_approximate_methods: true. |
| evaluator_retired | An evaluator behind the metric was retired. |
INSUFFICIENT_EVIDENCE
| Code | Meaning |
|---|---|
| no_observations | No case was observed for this metric. |
| missingness_exceeds_policy | More of the eligible cases are missing than the rule's max_missing_fraction allows. |
| missingness_unbounded | The method drops missing cases rather than bounding them, and the rule declares no max_missing_fraction. |
| evaluator_not_validated | A model judge behind the metric has not been validated against human labels, and require_validated_evaluators is on (the default). |
| evaluator_recalibration_required | The judge was validated on a served model this run's verdicts did not come from. |
| interval_unavailable | The metric has no interval to read. |
| insufficient_clusters | Fewer clusters than the policy's min_clusters. |
| interval_monte_carlo_uncertain | The threshold falls inside the simulation uncertainty of a clustered bound. |
| interval_overlaps_threshold | The interval contains the threshold. More cases would narrow it. |
| interval_unbounded | The interval has no bound on the side the rule reads. |
| missing_could_change_outcome | An observed_count rule: the missing cases could take the failures past max_failures. |
| interval_overlaps_zero, interval_overlaps_margin, interval_overlaps_margins | A comparison whose difference interval straddles zero or a margin. |
| insufficient_support | A slice comparison rule whose slice has fewer cases than its min_support. |
| family_correction_withheld | A rule in a families: entry that the Holm correction stopped before it. |
PASS and FAIL
| Code | State |
|---|---|
| lower_bound_meets_minimum, upper_bound_meets_maximum | PASS |
| upper_bound_below_minimum, lower_bound_above_maximum | FAIL |
| observed_failures_within_limit | PASS |
| observed_failures_exceed_limit | FAIL |
| difference_above_zero, lower_bound_above_margin, interval_within_margins | PASS (comparison) |
| difference_below_zero, upper_bound_below_margin, interval_outside_margins | FAIL (comparison) |
| cost_ceiling_exceeded | Said beside the state of a cost rule whose declared ceiling a recorded execution exceeded |
Hosted workspace only
A workspace that decides a pushed run itself can withhold a decision with ppi_not_verified (it did not verify the interval the decision rests on), execution_not_verified (the outputs did not come from a registered runner) or workspace_cannot_decide (it holds no copy of the policy, or could not read the evidence). See Gating.
How each case enters the denominator
Every metric reports four counts: n_total (cases in the suite), n_eligible, n_observed and n_missing, with n_eligible = n_observed + n_missing. Cases outside n_eligible are listed under exclusions with a reason.
| What happened to the case | Counts as | In the denominator |
|---|---|---|
| The evaluator returned pass or fail | observed, success or failure | yes |
| The evaluator declared itself not applicable (for example, no expected value to compare) | excluded, with the reason | no |
| The system call errored or timed out | missing, or failure under on_execution_error: fail | yes |
| The evaluator raised, or a judge's reply could not be read | missing | yes |
| The case never ran because the run was interrupted | missing, and the run is PARTIAL | yes |
A missing case is bounded, not dropped. For a pass rate the interval's lower bound treats every missing case as a failure and its upper bound as a success, so a run with many missing cases has a wide interval that cannot pass a demanding rule; a bounded mean substitutes the ends of its declared range the same way. A method that cannot bound missing cases (a ranking statistic, for example) drops them and records the assumption, and a rule over it reads missingness_unbounded until it declares max_missing_fraction.
An observed_count rule counts failures over the executed suite and reads no interval. It passes only when the observed failures plus every missing case still fit within max_failures.
Metrics with no admitted interval
A rule decides only on an interval whose method has been admitted by audit. Where none exists, the metric is still computed and shown, and a rule over it does not borrow an unvalidated method:
| Situation | What a rule over it reads |
|---|---|
| A score (mean) metric with no declared range, such as a custom score evaluator without score_range | MANUAL_REVIEW, unsupported_method |
| A mean, quantile, ranking or cost metric on a suite that declares group_id | MANUAL_REVIEW, unsupported_dependence_structure |
| A pass rate on a clustered suite | an approximate interval: MANUAL_REVIEW unless allow_approximate_methods: true, then the cluster checks above |
| A quantile or ranking metric with replicates above 1 | MANUAL_REVIEW, unsupported_method |
| Any metric on a suite with both group_id and replicates | MANUAL_REVIEW, unsupported_dependence_structure |
| A comparison on a clustered suite | MANUAL_REVIEW |
| A slice below min_slice_support | no interval, but slices never reach the gate |
| human_score, human_preference or cost_per_accepted metrics | refused when the file is read, exit 2 |
Exit codes
oloproof gate, oloproof run with a policy, and the other commands that decide all use the same codes.
| Code | Meaning |
|---|---|
| 0 | Nothing the policy blocks on: every rule passed, or those that did not are outside block_on. |
| 1 | A rule in block_on failed. |
| 2 | The configuration or invocation was wrong, or a system broke its contract; nothing was decided. |
| 3 | A rule in block_on read INSUFFICIENT_EVIDENCE. |
| 4 | A rule in block_on read MANUAL_REVIEW. |
| 5 | The run did not complete and block_on_partial_run is on (the default). |
When several apply, the code reported is the first of 1, 5, 4, 3. A state left out of block_on cannot change the exit code: with block_on: [FAIL] and warn_on: [INSUFFICIENT_EVIDENCE], an undecided rule warns and the gate exits 0. Exit 0 therefore means only that nothing your policy blocks on occurred, not that every rule passed. See Gating.
Where the work runs and where the data goes
Local, the default
oloproof run, oloproof gate and the SDK run on your machine. Every record (cases, outputs, artifacts, judgments, metrics and decisions) is written to .oloproof/store.sqlite beside oloproof.yaml, or under OLOPROOF_HOME when set. Nothing is sent to Oloproof. The only network traffic is what your configuration causes: calls to your HTTP system's URL, and calls a model judge or model classifier makes to its provider, which receive the case content they judge and bill you for it.
Pushing to a hosted workspace
oloproof push sends a run's evidence to the workspace you connected with oloproof login. By default it sends metrics, intervals, decisions and aggregate slices, and every record's identity, status, timings and usage, but not its content. Raw content is redacted field by field before anything leaves the machine, and a redacted record says which categories were withheld. A category goes only when egress: in oloproof.yaml lists it:
| Category | What it covers |
|---|---|
| raw_inputs | Scenario inputs, expected values and case metadata: the dataset rows |
| raw_outputs | What the system under test returned for each case |
| judge_rationales | The text a judge wrote explaining a verdict, which quotes the output |
| artifacts | Retrieval context, citations and trajectories recorded during a run |
| error_detail | Exception messages and details, which often carry the input verbatim |
| system_config | The declared configuration of the system under test and of its evaluators |
| label_notes | The note a person wrote beside a label, which often quotes the output |
| span_names | Trace, span, tool and agent names an instrumentation recorded |
Record digests are not recomputed after redaction, so a hosted record still names the original evidence, which stays on your machine. Redaction is not encryption, and a metric over a very small slice can still identify the cases behind it.
When reviewers label cases in the hosted review queue, their browser fetches case content from oloproof collect running on your side; it does not pass through the workspace. Provider keys a workspace uses are stored with oloproof credentials set, and oloproof credentials list shows their names, never their values. A managed job runs on a worker Oloproof operates, which does not run your Python code; oloproof job reports its result and exits on its gate.
In a hosted workspace the engine never reuses cached executions, judgments or analyses, because a push can write those caches; it recomputes them.