Skip to content

Guides

Configuration reference

Every field of oloproof.yaml and release.yaml, with its type, its default, the values it accepts and an example, taken from the models that read the files. Use it to look up a field; read the quickstart and gating pages to learn the workflow.

Both files are validated before anything runs. An unknown field, a misspelled field or a value of the wrong type is a configuration error and the command exits 2 without executing a case. Both files have JSON Schemas, which an editor that reads JSON Schema can use for completion. The installed package writes them, with the result schemas, into schemas/v1/ under the current directory: python -m oloproof_core.models.schema_export (the two are project_config.schema.json and release_policy.schema.json).

In the tables below, "required" means the file is refused without the field; any other field shows the value used when it is left out.

oloproof.yaml at a glance

A small, complete project. It runs a Python function locally, needs no network and no keys, and is the shape oloproof init scaffolds.

# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
  name: support-bot
  callable: app.bot:answer
evaluators:
  - type: exact_match
    criterion: correct_label
    field: label

Top-level fields

FieldTypeDefaultWhat it is
version11File format version. Only 1 exists.
projectstringrequiredThe project's name, shown in reports and used when pushing.
datasetpathrequiredThe suite file, JSONL, relative to the project. Its rows are described in Suites.
systemmappingrequiredThe system under test. See below.
concurrencymappingsystem: 8, judge: 4How many system calls and judge calls run at once.
evaluatorslistrequired, at least oneWhat is measured on each case. Each entry has a type.
metricslistemptyExtra metrics beyond the one each evaluator criterion already is.
predictivemappingabsentWhere a classifier's label, score and truth live. See Predictive models.
sliceslist of stringsemptyExploratory slices: metadata.<key>, relevant_position or context_truncated. They never reach the gate. See Slices.
min_slice_supportinteger, at least 130Below this many eligible cases a slice shows its estimate but no interval.
replicatesinteger, at least 11Measure every case this many times. The case stays the unit: replicates are aggregated within it before any interval is computed.
pricinglistemptyWhat you pay per million tokens, by model. Without it cost is reported in tokens and never in dollars.
egresslist of stringsemptyWhich raw content oloproof push may send to a hosted workspace. See Results and execution.

concurrency

FieldTypeDefault
systeminteger, at least 18
judgeinteger, at least 14

pricing entries

Oloproof ships no price table. Each entry names a model exactly as an evaluator's model: names it.

FieldTypeDefault
modelstringrequired
input_per_mtoknumber, 0 or morerequired
output_per_mtoknumber, 0 or morerequired
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
  name: support-bot
  callable: app.bot:answer
evaluators:
  - type: exact_match
    criterion: correct_label
    field: label
pricing:
  - model: my-judge-model
    input_per_mtok: 0.15
    output_per_mtok: 0.6
egress: [raw_outputs]

system

A system needs exactly one of callable, http or rag.

FieldTypeDefaultWhat it is
namestringrequiredThe system's name. Part of its version identity.
versionstringabsentYour label for this version. Required for an HTTP system. Part of its identity, so changing it invalidates cached executions.
callablemodule:attributeabsentA Python function, sync or async. It receives the case input and returns the output.
httpmappingabsentAn endpoint called once per case. See below.
ragmappingabsentA staged RAG class declared with @rag_system. See below.
configmappingemptyFree-form settings recorded with the system version. Changing them changes the version.
code_pathslist of glob patternsemptySource files whose contents enter a callable system's version. Without it only the callable's own module is hashed.
timeout_snumber above 0120Time limit per call for a callable system. An HTTP system uses http.timeout_s instead.
recordslist of artifact kindsemptyArtifact kinds a callable system records, such as retrieval/v1. Refused on an HTTP or RAG system.

system.http

FieldTypeDefaultWhat it is
urlstringrequiredWhere each case is sent.
methodGET, POST or PUTPOSTThe HTTP method.
output_pathdotted pathabsentWhich field of the JSON response is the output, such as result.answer. Absent means the whole body.
artifactsmapping of kind to dotted pathemptyResponse fields recorded as artifacts, such as retrieval/v1: debug.retrieval.
versionstringabsentUsed as the system's version when system.version is absent. An HTTP system needs one of the two.
timeout_snumber above 030Time limit per request.
# oloproof.yaml
version: 1
project: support-api
dataset: datasets/support.jsonl
system:
  name: support-api
  version: "2026-10-08"
  http:
    url: http://localhost:8000/answer
    output_path: answer
    artifacts:
      retrieval/v1: debug.retrieval
evaluators:
  - type: hit_rate
    k: 5

The request and response contract, and what happens on timeouts and HTTP errors, are in Results and execution.

system.rag

FieldTypeDefaultWhat it is
objectmodule:attributerequiredThe class declared with @rag_system, or an instance of it.
depthinteger, at least 1the class'sHow many passages retrieval returns.
top_kinteger, at least 1the class'sHow many of them reach generation.
token_budgetinteger, at least 1the class'sA token limit on the context. Needs the class's count_tokens(passage).
index_versionstringthe class'sPart of the retrieval identity. Change it whenever the index is rebuilt.

Settings given here override the ones the class declares. A staged system records its own retrieval/v1, context/v1 and citations/v1 artifacts, so records is refused beside it. See RAG.

evaluators

Every entry takes a type and these two common fields:

FieldTypeDefaultWhat it is
criterionstringrequired unless the type has a defaultThe name of what is measured. Each criterion is a metric, and a rule's metric: names it.
on_execution_errormissing or failmissingWhat a case whose system call failed counts as for this criterion. missing keeps it in the denominator as unobserved; fail counts it as a failure.

fail applies only to pass/fail evaluators; a score evaluator with it is a configuration error. on_execution_error is a YAML field; the SDK evaluator classes take no such argument, and an errored case counts as missing.

Evaluator types

"Reads" lists what the evaluator's verdict depends on, which is also what its cached judgment is keyed on. "SDK" names the class in oloproof.evaluators.

YAML typeReadsSDKNeeds network or a key
exact_matchoutput, expectedExactMatchno
containsoutput, expectedContainsno
regexoutputRegexno
json_schemaoutputJsonSchemano
rubric_judgeinput, output, expectedRubricJudgeyes, a model provider
model_classifieroutput (or the field named by text), optionally premiseYAML onlyyes, a TEI-compatible server
probability_judgethe case and outputYAML onlyyes, an OpenAI-compatible provider returning log probabilities
cascadeas its two stagesYAML onlyyes
hit_rate, recall, mrr, ndcgartifacts.retrieval, expectedHitRate, Recall, MRR, NDCGno
citation_validityartifacts.citations, artifacts.contextCitationValidityno
groundedness_judgeinput, output, artifacts.contextGroundednessyes
citation_support_judgeinput, output, artifacts.context, artifacts.citationsCitationSupportyes
agent_max_stepsartifacts.agent_trajectoryAgentMaxStepsno
agent_tool_calledartifacts.agent_trajectoryAgentToolCalledno
agent_no_tool_loopartifacts.agent_trajectoryAgentNoToolLoopno
agent_tool_sequenceartifacts.agent_trajectory, expectedAgentToolSequenceno
agent_no_undeclared_toolartifacts.agent_trajectory, expectedAgentNoUndeclaredToolno
agent_constraints_satisfiedartifacts.agent_trajectoryAgentConstraintsSatisfiedno
agent_routeartifacts.agent_trajectoryAgentRouteno
agent_tool_permissionsartifacts.agent_trajectoryAgentToolPermissionsno
agent_max_handoffsartifacts.agent_trajectoryAgentMaxHandoffsno
predictive_correctthe label field of output and expectedPredictiveCorrectno
predictive_recallas abovePredictiveRecallno
predictive_precisionas abovePredictivePrecisionno
predictive_absolute_erroras above, numericAbsoluteErrorno
predictive_brierthe score field of output, the label of expectedBrierno
predictive_log_lossas aboveLogLossno
predictive_rankingas abovePredictiveRankingno
no YAML typeartifacts.conversationConversationCompleted (SDK only)no
no YAML typeexpected, artifacts.conversationConversationJudge (SDK only)yes
no YAML typewhat you declare@evaluator and CustomEvaluator (SDK only)yours

A judge that calls a hosted model sends case content to that provider and is billed by it. Keys are read from the environment variable named in api_key_env; Oloproof never stores them in these files.

Deterministic evaluators

TypeFieldTypeDefault
exact_matchfielddotted path in the outputabsent: the whole output
exact_matchexpected_fielddotted path in expectedabsent: same as field
exact_matchstripbooleantrue
exact_matchcasefoldbooleanfalse
containsfield, expected_fieldas exact_matchabsent
regexpatternregular expressionrequired
regexfielddotted pathabsent
regexpass_ifmatch or no_matchmatch
json_schemaschemaa JSON Schema inline, or a path to a JSON file relative to the projectrequired
json_schemafielddotted pathabsent

Model judges

rubric_judge, groundedness_judge and citation_support_judge share these fields. Exactly one of rubric_file or rubric_text is required by rubric_judge; the two RAG judges take at most one and otherwise use a built-in rubric. Their criterion defaults to groundedness and citation_support.

FieldTypeDefault
provideranthropic, openai or openai_compatiblerequired
modelstringrequired
rubric_filepathabsent
rubric_textstringabsent
api_key_envenvironment variable nameANTHROPIC_API_KEY or OPENAI_API_KEY
base_urlURLthe provider's
temperaturenumber0
max_tokensinteger, at least 1512
timeout_snumber above 060

probability_judge asks a typed question and reads the model's probabilities:

FieldTypeDefault
provideropenai or openai_compatiblerequired
modelstringrequired
questionstringrequired
formyes_no, choice or scorerequired
min_probabilitynumber in (0, 1]required
optionsmapping of answer to descriptionfor choice
pass_optionslist of answersfor choice
levelsmapping of level to description, lowest firstfor score
pass_at_leasta levelfor score
calibrationslope (above 0), intercept, from_versionabsent
api_key_env, base_urlas aboveabsent
timeout_snumber above 060

cascade runs a cheap judge first and escalates the uncertain cases:

FieldTypeDefault
firsta probability_judge entryrequired
thena rubric_judge or probability_judge entryrequired
escalate_betweentwo probabilitiesrequired

The stages judge the cascade's own criterion; a stage naming a different one is refused.

model_classifier scores text with a trained model on a TEI-compatible server:

FieldTypeDefault
modelstringrequired
base_urlURLrequired
labelthe classifier label to readrequired
min_score or max_scorenumber in [0, 1], exactly onerequired
textwhich field is classifiedoutput
premisea second text, for pair classifiersabsent
api_key_envenvironment variable nameabsent
timeout_snumber above 030

RAG evaluators

TypeFieldTypeDefault
hit_rate, recall, mrr, ndcgkinteger, at least 15 for hit_rate and recall, 10 for mrr and ndcg
hit_rate, recall, mrr, ndcgrelevance_unitdoc or chunkdoc
hit_rate, recall, mrr, ndcgcriterionstring<type>_at_<k>, such as hit_rate_at_5
citation_validityrequire_citationsbooleanfalse
citation_validitycriterionstringcitations_valid

Agent evaluators

TypeFieldTypeDefault
agent_max_stepsmax_stepsinteger, at least 1required
agent_tool_calledtool_namestringrequired
agent_tool_calledmin_callsinteger, at least 11
agent_no_tool_loopmax_repeatsinteger, at least 12
agent_tool_sequenceorderedbooleantrue
agent_constraints_satisfiedconstraintslist of constraint namesempty
agent_tool_permissionspermissionsmapping of agent to allowed toolsrequired
agent_max_handoffsmax_handoffsinteger, 0 or morerequired

Each agent type has a default criterion, so it may be left out: its own type name, or one built from its setting (agent_steps_le_8, agent_tool_lookup_called, agent_handoffs_le_2). See Agents.

Predictive evaluators

TypeFieldTypeDefault
predictive_correct, predictive_recall, predictive_precisionpositiveany JSON valuetrue, or the predictive: block's
samefieldoutput fieldlabel, or predictive.label_field
sameexpected_fieldexpected fieldlabel, or predictive.expected_field
predictive_absolute_errortarget_rangetwo numbersrequired
predictive_absolute_errorfield, expected_fieldas abovelabel
predictive_brier, predictive_log_loss, predictive_rankingpositiveany JSON valuetrue, or the block's
samefieldoutput fieldscore, or predictive.score_field
sameexpected_fieldexpected fieldlabel, or the block's
predictive_log_lossclipnumber in (0, 0.5)required

A predictive evaluator that leaves positive, field or expected_field unwritten takes it from the predictive: block; a value it writes is kept.

predictive

FieldTypeDefault
label_fieldstringlabel
score_fieldstringscore
expected_fieldstringlabel
positiveany JSON valuetrue
calibration_binsinteger, at least 110
thresholdslist of numbersempty
averagemacro or microabsent: no aggregate

metrics

Each evaluator criterion is already a metric. A metrics: entry adds one more, discriminated by type.

typeFieldsWhat it is
quantileid, source, quantile in (0, 1)A quantile of latency_ms, input_tokens, output_tokens, cost_usd, agent_steps or agent_tool_calls.
rankingid, criterion, statistic: roc_auc or average_precisionA statistic over the order of a ranking criterion's scores.
human_score, human_preferenceidRefused: no admitted method reads these labels yet.
cost_per_acceptedid, criterion, cost_ceiling_usd, cost_ceiling_sourceRefused until its wiring is admitted by audit.
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
  name: support-bot
  callable: app.bot:answer
evaluators:
  - type: exact_match
    criterion: correct_label
    field: label
metrics:
  - id: latency_p95
    type: quantile
    source: latency_ms
    quantile: 0.95

release.yaml

The release policy: which rules decide, and which decisions block. Omitted settings keep their defaults, so a policy that names only its rules still blocks on FAIL, INSUFFICIENT_EVIDENCE and MANUAL_REVIEW.

# release.yaml
version: 1
rules:
  - id: label_accuracy
    metric: correct_label
    min: 0.8
FieldTypeDefaultWhat it is
version11File format version.
confidence_levelprobability0.95The level of every interval a rule reads.
block_onlist of decision statesFAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEWStates that make the gate block, and set the exit code.
warn_onlist of decision statesemptyStates that warn without blocking. Must not overlap block_on.
block_on_partial_runbooleantrueWhether a run that did not complete blocks, with exit 5.
require_validated_evaluatorsbooleantrueWhether a rule over a model judge withholds its decision until the judge is validated against human labels. Deterministic evaluators are exempt.
minimum_evaluator_agreementnumber in [0, 1]absentThe agreement with human labels a judge must reach, by its lower bound, before it may be validated.
maximum_evaluator_biasnumber in (0, 1]absentHow far a judge's pass rate may sit from the people's before it may be validated.
allow_approximate_methodsbooleanfalseWhether a rule may decide on an interval the engine marks approximate (the clustered binary interval). Otherwise it reads MANUAL_REVIEW.
min_clustersinteger, at least 1020Fewer clusters than this and a clustered rule reads INSUFFICIENT_EVIDENCE.
difference_methodbounded_paired_difference@1 or conditional_exact_paired_difference@1absent: the firstWhich admitted method bounds a paired binary-rate difference.
early_stoppingbooleanfalseRun cases in batches and stop once every rule is decided. See Gating.
early_stopping_seedinteger, 0 or moreabsentThe seed of the case order.
early_stopping_batch_sizeinteger, at least 125Cases per batch.
ruleslistrequired, at least oneThe rules. See below.
familieslistemptyRules whose false FAILs are controlled together.
review_rulemappingabsentRefused: the wiring is not admitted yet.

rules

One list holds both kinds. A run rule takes exactly one of min, max or max_failures. A comparison rule names its kind and decides a difference between two runs; see Comparison rules.

FieldTypeDefaultApplies to
idstringrequiredall
metrica metric id or criterionrequiredall
kindinterval_threshold, observed_count, superiority, non_inferiority, equivalenceinferred for run rulesall
minnumberabsentrun rules: PASS when the interval's lower bound is at least this
maxnumberabsentrun rules: PASS when the interval's upper bound is at most this
max_failuresinteger, 0 or moreabsentobserved_count: a count over the executed suite, no interval
marginnumber above 0, in the metric's unitsabsentnon_inferiority and equivalence; refused on superiority
directionmin or maxminnon_inferiority only: whether higher or lower is better
max_missing_fractionnumber in [0, 1]absentinterval and comparison rules
requires_manual_reviewbooleanfalseall: the rule always reads MANUAL_REVIEW
scopeglobal or a sliceglobalinterval and comparison rules
min_supportinteger, at least 1absentcomparison rules on a slice

families

FieldTypeDefault
idstringrequired
correctionholmholm
ruleslist of rule idsrequired, at least one
# release.yaml
version: 1
warn_on: [INSUFFICIENT_EVIDENCE]
block_on: [FAIL, MANUAL_REVIEW]
rules:
  - id: label_accuracy
    metric: correct_label
    min: 0.8
    max_missing_fraction: 0.05
  - id: no_regression
    metric: correct_label
    kind: non_inferiority
    margin: 0.02

Artifact kinds

An artifact is a typed record a system writes beside its output, such as what it retrieved. A kind is a lowercase name with an optional version, matching ^[a-z][a-z0-9_]*(/v[1-9][0-9]*)?$. Evaluators that need an artifact name it, and a run whose system does not declare a required kind is refused before it starts, rather than counting every case as missing.

KindWritten byRequired by
retrieval/v1current_case().retrieval(...), a @rag_system, or http.artifactshit_rate, recall, mrr, ndcg
context/v1current_case().context(...) or a @rag_systemcitation_validity, groundedness_judge, citation_support_judge
citations/v1current_case().citations(...) or a @rag_systemcitation_validity, citation_support_judge
agent_trajectory/v1current_case().agent_trajectory(...)every agent_* evaluator, and the agent_steps and agent_tool_calls sources
conversation/v1current_case().artifact(CONVERSATION, ...)ConversationCompleted, ConversationJudge
stage_timings/v1a @rag_systemnone; shown beside latency

A callable system declares the kinds it records in records: (or @system(records=...)); an HTTP system in http.artifacts; a staged RAG system records its own.

Versions, cache keys and invalidation

Oloproof reuses work whose inputs have not changed, and decides what "unchanged" means from content digests. Each is computed by the engine and recorded with the run.

RecordReused when these are identical
System versionname, version, config, and a code digest: a callable's module source (or every file matched by code_paths), an HTTP system's url, method, output_path and artifacts
Executionthe system version, the case's input and the replicate index. Only successful executions are reused.
Judgmentthe evaluator version (its type and every setting) and a digest of each field it reads, as listed in the evaluator table
Analysisthe analysis plan, the metric, the confidence level, the suite digest and every input it counted
Gateevery analysis, the policy digest, whether the run completed, and the effective status of each evaluator the decisions cite

What Oloproof cannot see is yours to declare:

  • An HTTP system's behaviour lives on the server. Change system.version whenever what is behind the URL changes, or an old cached output will stand for the new system.
  • A callable's helper modules are hashed only when code_paths matches them. Without it, editing a helper does not change the version.
  • A method or a callable object must declare a version, and the version must change when the object's state changes.
  • A RAG index is identified by index_version; change it when the index is rebuilt.
  • A model judge's identity is its settings, not the provider's weights. A provider updating the model behind the same name is not detected by the cache.
  • A custom @evaluator hashes the module file that defines it, and its judgments are reused across runs only when it declares cacheable=True. Built-in rubric judges are cacheable; deterministic evaluators are recomputed, which is cheap.

Cached work lives in the project's local store, .oloproof/store.sqlite beside oloproof.yaml (or under OLOPROOF_HOME). Deleting the store discards every cache and every run. In a hosted workspace the engine does not reuse cached executions, judgments or analyses, because a push can write them; it recomputes.