Guides
Tutorial: a classifier or structured output
Evaluate a Python function that labels support questions, read why the release is blocked, fix the misses, and compare the fix with the original, all on your own machine with no account, no network and no model.
What you will build
A support bot that returns a JSON object with an answer and a label (refund, account or other). You will hold it to three requirements: the label is right often enough, the output always has the right shape, and no answer leaks anything that looks like a US social security number. Two of those are format checks that need no reference answer; one measures task success against a reference label. The difference matters, and this page keeps them apart.
The terms used below (case, run, metric, interval, rule, gate) are defined in Concepts.
Prerequisites
- Python 3.11 or later.
- Oloproof, installed into a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof- The example project and the candidate change, which ship with the package. Copy both into new directories and work in the first; every file is also listed below, so you can type them instead:
oloproof init --example support_bot support-classifier
oloproof init --example classification support-change
cd support-classifierNo API key, provider account or network access is used anywhere on this page.
The files
support-classifier/
app.py the application under test (a Python callable)
oloproof.yaml the suite: dataset, system, evaluators
release.yaml the release policy: rules the run is decided against
data/support.jsonl 18 cases, one JSON object per line
rubrics/helpful.md a judge rubric, unused hereRun every command from the support-classifier/ directory. Oloproof keeps its store in .oloproof/ there; delete that directory to start again from nothing.
The application and its adapter
Your application is reached through an adapter. For a Python application the adapter is the function itself: Oloproof imports it, calls it once per case with the case's input, and records the dictionary it returns as that case's output.
# app.py
from typing import Any
from oloproof import system
@system(name="support-bot", version="slice-a-example")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if "refund" in question:
return {"answer": "Refunds are available within 30 days when the order is eligible.",
"label": "refund"}
if "password" in question or "login" in question:
return {"answer": "Use password reset, then contact support if the login still fails.",
"label": "account"}
return {"answer": "A support specialist will follow up with the next step.", "label": "other"}To evaluate your own classifier, keep its code where it is and write a thin function like this one that calls it and returns a dictionary. The function may be async. Oloproof calls it; it does not host, sandbox or reset your application, so any state your application keeps between calls is yours to manage.
oloproof.yaml names that function and the evaluators:
version: 1
project: support-bot-example
dataset: data/support.jsonl
system:
name: support-bot
version: slice-a-example
callable: app:answer
timeout_s: 30
evaluators:
- type: exact_match
criterion: exact_label
field: label
- type: json_schema
criterion: format_valid
field: null
schema:
type: object
required: [answer, label]
properties:
answer: {type: string}
label: {type: string}
additionalProperties: false
- type: regex
criterion: pii_free
field: answer
pattern: '\b\d{3}-\d{2}-\d{4}\b'
pass_if: no_matchOutputs are cached on the function's source, the declared version and config. If the function reads other files (a prompt, a rules table), list them under system.code_paths, so editing them runs the system again.
The dataset
One case per line. input is exactly what your function receives as case; expected is the reference the exact_match evaluator compares with:
{"id":"refund_00","input":{"question":"Can I get a refund for yesterday's order?"},"expected":{"label":"refund"}}
{"id":"account_04","input":{"question":"I can't sign in on my new phone."},"expected":{"label":"account"}}
{"id":"other_04","input":{"question":"I don't want a refund, I just need a copy of my receipt."},"expected":{"label":"other"}}The function returns, for each case, an object such as {"answer": "Use password reset, ...", "label": "account"}.
Choosing the evaluators
| Criterion | Evaluator | Needs expected | What it measures |
|---|---|---|---|
| exact_label | exact_match on label | yes | task success: the label is the right one |
| format_valid | json_schema over the whole output | no | format: the object has exactly the two string fields |
| pii_free | regex on answer, pass_if: no_match | no | a safety property of the text |
A format check passes a well-formed wrong answer, so it can never stand in for task success. A task check needs a reference for every case; where a case has none, exact_match cannot score it. Deterministic evaluators need no validation against people: running one twice gives the same verdict.
The policy
release.yaml is what the run is decided against:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
rules:
- id: exact-label-floor
metric: exact_label
min: 0.70
- id: valid-format
metric: format_valid
kind: observed_count
max_failures: 0
- id: pii-free
metric: pii_free
kind: observed_count
max_failures: 0exact-label-floor says the label must be right at least 70% of the time, and passes only when the whole 95% interval is at or above 0.70. The two observed_count rules allow no failure at all on the cases you ran; they describe these cases, not every question users will ask.
Run it
oloproof runReal output, trimmed:
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ exact-label-floor │ exact_label │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ valid-format │ format_valid │ PASS │ observed_failures_within_limit │
│ pii-free │ pii_free │ PASS │ observed_failures_within_limit │
│ exact_label │ 72.2% │ [46.5%, 90.4%] │ 13 / 18 observed · 0 missing · 0 excluded │
│ format_valid │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
│ pii_free │ 100.0% │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/18 miss; judgment 0 hit/54 missHow to read it:
- 13 of 18 labels are right, 72.2%. That is above 0.70, but the interval reaches down to 46.5%: 18 cases cannot show the true rate is at least 0.70. So the rule is INSUFFICIENT_EVIDENCE, not PASS and not FAIL.
- Every output has the right shape and none contains an SSN-like number, so both format rules pass.
- block_on lists INSUFFICIENT_EVIDENCE, so the gate blocks and the command exits 3. Exit 0 would mean nothing the policy blocks on; Gating CI lists every code.
Run it again and the cache line reads execution 18 hit/0 miss: nothing changed, so the function is not called.
Inspect the failures
oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_04RUN_ID is the id on the first line of the run's output.
5 of 18 cases failed, errored or did not finish
refund_04
output: {"answer": "A support specialist will follow up with the next step.", "label": "other"}
exact_label: failed
...
other_04
output: {"answer": "Refunds are available within 30 days when the order is eligible.", "label": "refund"}
exact_label: failedcase refund_04
input: {
"question": "I was charged twice this month and want my money back."
}
expected: {
"label": "refund"
}
execution: OK, 1 ms
output: {
"answer": "A support specialist will follow up with the next step.",
"label": "other"
}
judgments:
exact_label: failed
format_valid: passed
pii_free: passedThe pattern is plain once you read the inputs: "money back", "reverse the payment", "sign in" and "two-factor" are not in the keyword lists, and other_04 says "I don't want a refund", which the word "refund" matches anyway. Note that refund_04 passes both format checks while being wrong: that is the gap between checking format and measuring success.
Two next actions are meaningful here. Fix the misses (below), or add cases: with more cases at the same accuracy the interval narrows, and oloproof plan RUN_ID --run estimates how many.
Make a real change
Copy ../support-change/app.py over app.py. It adds the missed phrasings:
REFUND_WORDS = ("refund", "money back", "reverse the payment")
ACCOUNT_WORDS = ("password", "login", "sign in", "two-factor")
@system(name="support-bot", version="keywords-v2")
def answer(case: dict[str, Any]) -> dict[str, str]:
question = str(case["question"]).lower()
if any(word in question for word in REFUND_WORDS):
...and set version: keywords-v2 under system in oloproof.yaml, so the run is recorded as the new version. Then:
oloproof runGate: ALLOW (exit 0)
│ exact-label-floor │ exact_label │ PASS │ lower_bound_meets_minimum │
│ exact_label │ 94.4% │ [72.7%, 99.9%] │ 17 / 18 observed · 0 missing · 0 excluded │17 of 18 are right and the interval's lower bound, 72.7%, clears 0.70, so the rule passes and the command exits 0. other_04 still fails: the fix did not touch negation.
Compare the candidate with the baseline
The run rule asks whether the candidate meets your floor. A comparison asks how it differs from the baseline, case by case. Copy ../support-change/compare.yaml into the project; it holds one comparison rule:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: label-no-regression
kind: non_inferiority
metric: exact_label
margin: 0.10oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlComparison sha256:de76... of run_01M4...TJAD against run_01M4...ECVEF · 18 paired cases
exact_label: +22.2 points [-12.9, +57.0] · 18 paired · 0 missing · 0 excluded
format_valid: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
pii_free: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
Decisions
label-no-regression exact_label non-inferiority, margin 10.0 points INSUFFICIENT_EVIDENCE interval_overlaps_margin
about 3 more paired cases would decide it, if the difference holds (21 in total at 22% discordance)
Gate: BLOCK (exit 3)The candidate fixed four cases and broke none, an estimated gain of 22 points. But only four cases changed, and 18 paired cases leave an interval from 12.9 points worse to 57 points better, which crosses the 10-point margin. The comparison cannot yet rule out that the candidate is worse by more than you accept, so it is INSUFFICIENT_EVIDENCE and exits 3. The line beneath it is the sizing estimate. Without --policy, compare prints the differences, says the project's release.yaml declares no comparison rule, and exits 0 because nothing was decided.
Comparing a candidate to a baseline explains the margin and the other rule kinds.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| ModuleNotFoundError for app | Run from the directory holding app.py, or give callable a module path importable from there. |
| A rule names a metric no evaluator produces | The rule's metric must equal an evaluator's criterion; the error lists the metrics there are. |
| You edited the classifier and the run reused every output | The cache follows the callable's source; a helper file it reads must be listed under system.code_paths. |
| exact_label reports cases as missing | Those executions raised or timed out; oloproof inspect RUN_ID --failures shows each error. |
| The run exits 3 with a high estimate | The interval, not the estimate, decides. Add cases or accept a lower floor, decided before the run. |
Limitations
- The SDK and YAML report pass rates per evaluator. There is no confusion matrix or per-class precision and recall for a classifier like this one; a predictive: block does those for a scored model (Predictive models).
- observed_count rules describe the cases you ran; they make no claim about unseen inputs.
- A comparison over 18 cases resolves only large differences. Fifty or more real cases are a more useful floor.
- oloproof.yaml cannot name a custom @evaluator; that needs the SDK (The SDK).