Guides
Configuration reference
Every field of oloproof.yaml and release.yaml, with its type, its default, the values it accepts and an example, taken from the models that read the files. Use it to look up a field; read the quickstart and gating pages to learn the workflow.
Both files are validated before anything runs. An unknown field, a misspelled field or a value of the wrong type is a configuration error and the command exits 2 without executing a case. Both files have JSON Schemas, which an editor that reads JSON Schema can use for completion. The installed package writes them, with the result schemas, into schemas/v1/ under the current directory: python -m oloproof_core.models.schema_export (the two are project_config.schema.json and release_policy.schema.json).
In the tables below, "required" means the file is refused without the field; any other field shows the value used when it is left out.
oloproof.yaml at a glance
A small, complete project. It runs a Python function locally, needs no network and no keys, and is the shape oloproof init scaffolds.
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: labelTop-level fields
| Field | Type | Default | What it is |
|---|---|---|---|
| version | 1 | 1 | File format version. Only 1 exists. |
| project | string | required | The project's name, shown in reports and used when pushing. |
| dataset | path | required | The suite file, JSONL, relative to the project. Its rows are described in Suites. |
| system | mapping | required | The system under test. See below. |
| concurrency | mapping | system: 8, judge: 4 | How many system calls and judge calls run at once. |
| evaluators | list | required, at least one | What is measured on each case. Each entry has a type. |
| metrics | list | empty | Extra metrics beyond the one each evaluator criterion already is. |
| predictive | mapping | absent | Where a classifier's label, score and truth live. See Predictive models. |
| slices | list of strings | empty | Exploratory slices: metadata.<key>, relevant_position or context_truncated. They never reach the gate. See Slices. |
| min_slice_support | integer, at least 1 | 30 | Below this many eligible cases a slice shows its estimate but no interval. |
| replicates | integer, at least 1 | 1 | Measure every case this many times. The case stays the unit: replicates are aggregated within it before any interval is computed. |
| pricing | list | empty | What you pay per million tokens, by model. Without it cost is reported in tokens and never in dollars. |
| egress | list of strings | empty | Which raw content oloproof push may send to a hosted workspace. See Results and execution. |
concurrency
| Field | Type | Default |
|---|---|---|
| system | integer, at least 1 | 8 |
| judge | integer, at least 1 | 4 |
pricing entries
Oloproof ships no price table. Each entry names a model exactly as an evaluator's model: names it.
| Field | Type | Default |
|---|---|---|
| model | string | required |
| input_per_mtok | number, 0 or more | required |
| output_per_mtok | number, 0 or more | required |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
pricing:
- model: my-judge-model
input_per_mtok: 0.15
output_per_mtok: 0.6
egress: [raw_outputs]system
A system needs exactly one of callable, http or rag.
| Field | Type | Default | What it is |
|---|---|---|---|
| name | string | required | The system's name. Part of its version identity. |
| version | string | absent | Your label for this version. Required for an HTTP system. Part of its identity, so changing it invalidates cached executions. |
| callable | module:attribute | absent | A Python function, sync or async. It receives the case input and returns the output. |
| http | mapping | absent | An endpoint called once per case. See below. |
| rag | mapping | absent | A staged RAG class declared with @rag_system. See below. |
| config | mapping | empty | Free-form settings recorded with the system version. Changing them changes the version. |
| code_paths | list of glob patterns | empty | Source files whose contents enter a callable system's version. Without it only the callable's own module is hashed. |
| timeout_s | number above 0 | 120 | Time limit per call for a callable system. An HTTP system uses http.timeout_s instead. |
| records | list of artifact kinds | empty | Artifact kinds a callable system records, such as retrieval/v1. Refused on an HTTP or RAG system. |
system.http
| Field | Type | Default | What it is |
|---|---|---|---|
| url | string | required | Where each case is sent. |
| method | GET, POST or PUT | POST | The HTTP method. |
| output_path | dotted path | absent | Which field of the JSON response is the output, such as result.answer. Absent means the whole body. |
| artifacts | mapping of kind to dotted path | empty | Response fields recorded as artifacts, such as retrieval/v1: debug.retrieval. |
| version | string | absent | Used as the system's version when system.version is absent. An HTTP system needs one of the two. |
| timeout_s | number above 0 | 30 | Time limit per request. |
# oloproof.yaml
version: 1
project: support-api
dataset: datasets/support.jsonl
system:
name: support-api
version: "2026-10-08"
http:
url: http://localhost:8000/answer
output_path: answer
artifacts:
retrieval/v1: debug.retrieval
evaluators:
- type: hit_rate
k: 5The request and response contract, and what happens on timeouts and HTTP errors, are in Results and execution.
system.rag
| Field | Type | Default | What it is |
|---|---|---|---|
| object | module:attribute | required | The class declared with @rag_system, or an instance of it. |
| depth | integer, at least 1 | the class's | How many passages retrieval returns. |
| top_k | integer, at least 1 | the class's | How many of them reach generation. |
| token_budget | integer, at least 1 | the class's | A token limit on the context. Needs the class's count_tokens(passage). |
| index_version | string | the class's | Part of the retrieval identity. Change it whenever the index is rebuilt. |
Settings given here override the ones the class declares. A staged system records its own retrieval/v1, context/v1 and citations/v1 artifacts, so records is refused beside it. See RAG.
evaluators
Every entry takes a type and these two common fields:
| Field | Type | Default | What it is |
|---|---|---|---|
| criterion | string | required unless the type has a default | The name of what is measured. Each criterion is a metric, and a rule's metric: names it. |
| on_execution_error | missing or fail | missing | What a case whose system call failed counts as for this criterion. missing keeps it in the denominator as unobserved; fail counts it as a failure. |
fail applies only to pass/fail evaluators; a score evaluator with it is a configuration error. on_execution_error is a YAML field; the SDK evaluator classes take no such argument, and an errored case counts as missing.
Evaluator types
"Reads" lists what the evaluator's verdict depends on, which is also what its cached judgment is keyed on. "SDK" names the class in oloproof.evaluators.
| YAML type | Reads | SDK | Needs network or a key |
|---|---|---|---|
| exact_match | output, expected | ExactMatch | no |
| contains | output, expected | Contains | no |
| regex | output | Regex | no |
| json_schema | output | JsonSchema | no |
| rubric_judge | input, output, expected | RubricJudge | yes, a model provider |
| model_classifier | output (or the field named by text), optionally premise | YAML only | yes, a TEI-compatible server |
| probability_judge | the case and output | YAML only | yes, an OpenAI-compatible provider returning log probabilities |
| cascade | as its two stages | YAML only | yes |
| hit_rate, recall, mrr, ndcg | artifacts.retrieval, expected | HitRate, Recall, MRR, NDCG | no |
| citation_validity | artifacts.citations, artifacts.context | CitationValidity | no |
| groundedness_judge | input, output, artifacts.context | Groundedness | yes |
| citation_support_judge | input, output, artifacts.context, artifacts.citations | CitationSupport | yes |
| agent_max_steps | artifacts.agent_trajectory | AgentMaxSteps | no |
| agent_tool_called | artifacts.agent_trajectory | AgentToolCalled | no |
| agent_no_tool_loop | artifacts.agent_trajectory | AgentNoToolLoop | no |
| agent_tool_sequence | artifacts.agent_trajectory, expected | AgentToolSequence | no |
| agent_no_undeclared_tool | artifacts.agent_trajectory, expected | AgentNoUndeclaredTool | no |
| agent_constraints_satisfied | artifacts.agent_trajectory | AgentConstraintsSatisfied | no |
| agent_route | artifacts.agent_trajectory | AgentRoute | no |
| agent_tool_permissions | artifacts.agent_trajectory | AgentToolPermissions | no |
| agent_max_handoffs | artifacts.agent_trajectory | AgentMaxHandoffs | no |
| predictive_correct | the label field of output and expected | PredictiveCorrect | no |
| predictive_recall | as above | PredictiveRecall | no |
| predictive_precision | as above | PredictivePrecision | no |
| predictive_absolute_error | as above, numeric | AbsoluteError | no |
| predictive_brier | the score field of output, the label of expected | Brier | no |
| predictive_log_loss | as above | LogLoss | no |
| predictive_ranking | as above | PredictiveRanking | no |
| no YAML type | artifacts.conversation | ConversationCompleted (SDK only) | no |
| no YAML type | expected, artifacts.conversation | ConversationJudge (SDK only) | yes |
| no YAML type | what you declare | @evaluator and CustomEvaluator (SDK only) | yours |
A judge that calls a hosted model sends case content to that provider and is billed by it. Keys are read from the environment variable named in api_key_env; Oloproof never stores them in these files.
Deterministic evaluators
| Type | Field | Type | Default |
|---|---|---|---|
| exact_match | field | dotted path in the output | absent: the whole output |
| exact_match | expected_field | dotted path in expected | absent: same as field |
| exact_match | strip | boolean | true |
| exact_match | casefold | boolean | false |
| contains | field, expected_field | as exact_match | absent |
| regex | pattern | regular expression | required |
| regex | field | dotted path | absent |
| regex | pass_if | match or no_match | match |
| json_schema | schema | a JSON Schema inline, or a path to a JSON file relative to the project | required |
| json_schema | field | dotted path | absent |
Model judges
rubric_judge, groundedness_judge and citation_support_judge share these fields. Exactly one of rubric_file or rubric_text is required by rubric_judge; the two RAG judges take at most one and otherwise use a built-in rubric. Their criterion defaults to groundedness and citation_support.
| Field | Type | Default |
|---|---|---|
| provider | anthropic, openai or openai_compatible | required |
| model | string | required |
| rubric_file | path | absent |
| rubric_text | string | absent |
| api_key_env | environment variable name | ANTHROPIC_API_KEY or OPENAI_API_KEY |
| base_url | URL | the provider's |
| temperature | number | 0 |
| max_tokens | integer, at least 1 | 512 |
| timeout_s | number above 0 | 60 |
probability_judge asks a typed question and reads the model's probabilities:
| Field | Type | Default |
|---|---|---|
| provider | openai or openai_compatible | required |
| model | string | required |
| question | string | required |
| form | yes_no, choice or score | required |
| min_probability | number in (0, 1] | required |
| options | mapping of answer to description | for choice |
| pass_options | list of answers | for choice |
| levels | mapping of level to description, lowest first | for score |
| pass_at_least | a level | for score |
| calibration | slope (above 0), intercept, from_version | absent |
| api_key_env, base_url | as above | absent |
| timeout_s | number above 0 | 60 |
cascade runs a cheap judge first and escalates the uncertain cases:
| Field | Type | Default |
|---|---|---|
| first | a probability_judge entry | required |
| then | a rubric_judge or probability_judge entry | required |
| escalate_between | two probabilities | required |
The stages judge the cascade's own criterion; a stage naming a different one is refused.
model_classifier scores text with a trained model on a TEI-compatible server:
| Field | Type | Default |
|---|---|---|
| model | string | required |
| base_url | URL | required |
| label | the classifier label to read | required |
| min_score or max_score | number in [0, 1], exactly one | required |
| text | which field is classified | output |
| premise | a second text, for pair classifiers | absent |
| api_key_env | environment variable name | absent |
| timeout_s | number above 0 | 30 |
RAG evaluators
| Type | Field | Type | Default |
|---|---|---|---|
| hit_rate, recall, mrr, ndcg | k | integer, at least 1 | 5 for hit_rate and recall, 10 for mrr and ndcg |
| hit_rate, recall, mrr, ndcg | relevance_unit | doc or chunk | doc |
| hit_rate, recall, mrr, ndcg | criterion | string | <type>_at_<k>, such as hit_rate_at_5 |
| citation_validity | require_citations | boolean | false |
| citation_validity | criterion | string | citations_valid |
Agent evaluators
| Type | Field | Type | Default |
|---|---|---|---|
| agent_max_steps | max_steps | integer, at least 1 | required |
| agent_tool_called | tool_name | string | required |
| agent_tool_called | min_calls | integer, at least 1 | 1 |
| agent_no_tool_loop | max_repeats | integer, at least 1 | 2 |
| agent_tool_sequence | ordered | boolean | true |
| agent_constraints_satisfied | constraints | list of constraint names | empty |
| agent_tool_permissions | permissions | mapping of agent to allowed tools | required |
| agent_max_handoffs | max_handoffs | integer, 0 or more | required |
Each agent type has a default criterion, so it may be left out: its own type name, or one built from its setting (agent_steps_le_8, agent_tool_lookup_called, agent_handoffs_le_2). See Agents.
Predictive evaluators
| Type | Field | Type | Default |
|---|---|---|---|
| predictive_correct, predictive_recall, predictive_precision | positive | any JSON value | true, or the predictive: block's |
| same | field | output field | label, or predictive.label_field |
| same | expected_field | expected field | label, or predictive.expected_field |
| predictive_absolute_error | target_range | two numbers | required |
| predictive_absolute_error | field, expected_field | as above | label |
| predictive_brier, predictive_log_loss, predictive_ranking | positive | any JSON value | true, or the block's |
| same | field | output field | score, or predictive.score_field |
| same | expected_field | expected field | label, or the block's |
| predictive_log_loss | clip | number in (0, 0.5) | required |
A predictive evaluator that leaves positive, field or expected_field unwritten takes it from the predictive: block; a value it writes is kept.
predictive
| Field | Type | Default |
|---|---|---|
| label_field | string | label |
| score_field | string | score |
| expected_field | string | label |
| positive | any JSON value | true |
| calibration_bins | integer, at least 1 | 10 |
| thresholds | list of numbers | empty |
| average | macro or micro | absent: no aggregate |
metrics
Each evaluator criterion is already a metric. A metrics: entry adds one more, discriminated by type.
| type | Fields | What it is |
|---|---|---|
| quantile | id, source, quantile in (0, 1) | A quantile of latency_ms, input_tokens, output_tokens, cost_usd, agent_steps or agent_tool_calls. |
| ranking | id, criterion, statistic: roc_auc or average_precision | A statistic over the order of a ranking criterion's scores. |
| human_score, human_preference | id | Refused: no admitted method reads these labels yet. |
| cost_per_accepted | id, criterion, cost_ceiling_usd, cost_ceiling_source | Refused until its wiring is admitted by audit. |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
metrics:
- id: latency_p95
type: quantile
source: latency_ms
quantile: 0.95release.yaml
The release policy: which rules decide, and which decisions block. Omitted settings keep their defaults, so a policy that names only its rules still blocks on FAIL, INSUFFICIENT_EVIDENCE and MANUAL_REVIEW.
# release.yaml
version: 1
rules:
- id: label_accuracy
metric: correct_label
min: 0.8| Field | Type | Default | What it is |
|---|---|---|---|
| version | 1 | 1 | File format version. |
| confidence_level | probability | 0.95 | The level of every interval a rule reads. |
| block_on | list of decision states | FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW | States that make the gate block, and set the exit code. |
| warn_on | list of decision states | empty | States that warn without blocking. Must not overlap block_on. |
| block_on_partial_run | boolean | true | Whether a run that did not complete blocks, with exit 5. |
| require_validated_evaluators | boolean | true | Whether a rule over a model judge withholds its decision until the judge is validated against human labels. Deterministic evaluators are exempt. |
| minimum_evaluator_agreement | number in [0, 1] | absent | The agreement with human labels a judge must reach, by its lower bound, before it may be validated. |
| maximum_evaluator_bias | number in (0, 1] | absent | How far a judge's pass rate may sit from the people's before it may be validated. |
| allow_approximate_methods | boolean | false | Whether a rule may decide on an interval the engine marks approximate (the clustered binary interval). Otherwise it reads MANUAL_REVIEW. |
| min_clusters | integer, at least 10 | 20 | Fewer clusters than this and a clustered rule reads INSUFFICIENT_EVIDENCE. |
| difference_method | bounded_paired_difference@1 or conditional_exact_paired_difference@1 | absent: the first | Which admitted method bounds a paired binary-rate difference. |
| early_stopping | boolean | false | Run cases in batches and stop once every rule is decided. See Gating. |
| early_stopping_seed | integer, 0 or more | absent | The seed of the case order. |
| early_stopping_batch_size | integer, at least 1 | 25 | Cases per batch. |
| rules | list | required, at least one | The rules. See below. |
| families | list | empty | Rules whose false FAILs are controlled together. |
| review_rule | mapping | absent | Refused: the wiring is not admitted yet. |
rules
One list holds both kinds. A run rule takes exactly one of min, max or max_failures. A comparison rule names its kind and decides a difference between two runs; see Comparison rules.
| Field | Type | Default | Applies to |
|---|---|---|---|
| id | string | required | all |
| metric | a metric id or criterion | required | all |
| kind | interval_threshold, observed_count, superiority, non_inferiority, equivalence | inferred for run rules | all |
| min | number | absent | run rules: PASS when the interval's lower bound is at least this |
| max | number | absent | run rules: PASS when the interval's upper bound is at most this |
| max_failures | integer, 0 or more | absent | observed_count: a count over the executed suite, no interval |
| margin | number above 0, in the metric's units | absent | non_inferiority and equivalence; refused on superiority |
| direction | min or max | min | non_inferiority only: whether higher or lower is better |
| max_missing_fraction | number in [0, 1] | absent | interval and comparison rules |
| requires_manual_review | boolean | false | all: the rule always reads MANUAL_REVIEW |
| scope | global or a slice | global | interval and comparison rules |
| min_support | integer, at least 1 | absent | comparison rules on a slice |
families
| Field | Type | Default |
|---|---|---|
| id | string | required |
| correction | holm | holm |
| rules | list of rule ids | required, at least one |
# release.yaml
version: 1
warn_on: [INSUFFICIENT_EVIDENCE]
block_on: [FAIL, MANUAL_REVIEW]
rules:
- id: label_accuracy
metric: correct_label
min: 0.8
max_missing_fraction: 0.05
- id: no_regression
metric: correct_label
kind: non_inferiority
margin: 0.02Artifact kinds
An artifact is a typed record a system writes beside its output, such as what it retrieved. A kind is a lowercase name with an optional version, matching ^[a-z][a-z0-9_]*(/v[1-9][0-9]*)?$. Evaluators that need an artifact name it, and a run whose system does not declare a required kind is refused before it starts, rather than counting every case as missing.
| Kind | Written by | Required by |
|---|---|---|
| retrieval/v1 | current_case().retrieval(...), a @rag_system, or http.artifacts | hit_rate, recall, mrr, ndcg |
| context/v1 | current_case().context(...) or a @rag_system | citation_validity, groundedness_judge, citation_support_judge |
| citations/v1 | current_case().citations(...) or a @rag_system | citation_validity, citation_support_judge |
| agent_trajectory/v1 | current_case().agent_trajectory(...) | every agent_* evaluator, and the agent_steps and agent_tool_calls sources |
| conversation/v1 | current_case().artifact(CONVERSATION, ...) | ConversationCompleted, ConversationJudge |
| stage_timings/v1 | a @rag_system | none; shown beside latency |
A callable system declares the kinds it records in records: (or @system(records=...)); an HTTP system in http.artifacts; a staged RAG system records its own.
Versions, cache keys and invalidation
Oloproof reuses work whose inputs have not changed, and decides what "unchanged" means from content digests. Each is computed by the engine and recorded with the run.
| Record | Reused when these are identical |
|---|---|
| System version | name, version, config, and a code digest: a callable's module source (or every file matched by code_paths), an HTTP system's url, method, output_path and artifacts |
| Execution | the system version, the case's input and the replicate index. Only successful executions are reused. |
| Judgment | the evaluator version (its type and every setting) and a digest of each field it reads, as listed in the evaluator table |
| Analysis | the analysis plan, the metric, the confidence level, the suite digest and every input it counted |
| Gate | every analysis, the policy digest, whether the run completed, and the effective status of each evaluator the decisions cite |
What Oloproof cannot see is yours to declare:
- An HTTP system's behaviour lives on the server. Change system.version whenever what is behind the URL changes, or an old cached output will stand for the new system.
- A callable's helper modules are hashed only when code_paths matches them. Without it, editing a helper does not change the version.
- A method or a callable object must declare a version, and the version must change when the object's state changes.
- A RAG index is identified by index_version; change it when the index is rebuilt.
- A model judge's identity is its settings, not the provider's weights. A provider updating the model behind the same name is not detected by the cache.
- A custom @evaluator hashes the module file that defines it, and its judgments are reused across runs only when it declares cacheable=True. Built-in rubric judges are cacheable; deterministic evaluators are recomputed, which is cheap.
Cached work lives in the project's local store, .oloproof/store.sqlite beside oloproof.yaml (or under OLOPROOF_HOME). Deleting the store discards every cache and every run. In a hosted workspace the engine does not reuse cached executions, judgments or analyses, because a push can write them; it recomputes.