Skip to content

Guides

Tutorial: a classifier or structured output

Evaluate a Python function that labels support questions, read why the release is blocked, fix the misses, and compare the fix with the original, all on your own machine with no account, no network and no model.

What you will build

A support bot that returns a JSON object with an answer and a label (refund, account or other). You will hold it to three requirements: the label is right often enough, the output always has the right shape, and no answer leaks anything that looks like a US social security number. Two of those are format checks that need no reference answer; one measures task success against a reference label. The difference matters, and this page keeps them apart.

The terms used below (case, run, metric, interval, rule, gate) are defined in Concepts.

Prerequisites

  • Python 3.11 or later.
  • Oloproof, installed into a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • The example project and the candidate change, which ship with the package. Copy both into new directories and work in the first; every file is also listed below, so you can type them instead:
oloproof init --example support_bot support-classifier
oloproof init --example classification support-change
cd support-classifier

No API key, provider account or network access is used anywhere on this page.

The files

support-classifier/
  app.py              the application under test (a Python callable)
  oloproof.yaml       the suite: dataset, system, evaluators
  release.yaml        the release policy: rules the run is decided against
  data/support.jsonl  18 cases, one JSON object per line
  rubrics/helpful.md  a judge rubric, unused here

Run every command from the support-classifier/ directory. Oloproof keeps its store in .oloproof/ there; delete that directory to start again from nothing.

The application and its adapter

Your application is reached through an adapter. For a Python application the adapter is the function itself: Oloproof imports it, calls it once per case with the case's input, and records the dictionary it returns as that case's output.

# app.py
from typing import Any

from oloproof import system


@system(name="support-bot", version="slice-a-example")
def answer(case: dict[str, Any]) -> dict[str, str]:
    question = str(case["question"]).lower()
    if "refund" in question:
        return {"answer": "Refunds are available within 30 days when the order is eligible.",
                "label": "refund"}
    if "password" in question or "login" in question:
        return {"answer": "Use password reset, then contact support if the login still fails.",
                "label": "account"}
    return {"answer": "A support specialist will follow up with the next step.", "label": "other"}

To evaluate your own classifier, keep its code where it is and write a thin function like this one that calls it and returns a dictionary. The function may be async. Oloproof calls it; it does not host, sandbox or reset your application, so any state your application keeps between calls is yours to manage.

oloproof.yaml names that function and the evaluators:

version: 1
project: support-bot-example
dataset: data/support.jsonl
system:
  name: support-bot
  version: slice-a-example
  callable: app:answer
  timeout_s: 30
evaluators:
  - type: exact_match
    criterion: exact_label
    field: label
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [answer, label]
      properties:
        answer: {type: string}
        label: {type: string}
      additionalProperties: false
  - type: regex
    criterion: pii_free
    field: answer
    pattern: '\b\d{3}-\d{2}-\d{4}\b'
    pass_if: no_match

Outputs are cached on the function's source, the declared version and config. If the function reads other files (a prompt, a rules table), list them under system.code_paths, so editing them runs the system again.

The dataset

One case per line. input is exactly what your function receives as case; expected is the reference the exact_match evaluator compares with:

{"id":"refund_00","input":{"question":"Can I get a refund for yesterday's order?"},"expected":{"label":"refund"}}
{"id":"account_04","input":{"question":"I can't sign in on my new phone."},"expected":{"label":"account"}}
{"id":"other_04","input":{"question":"I don't want a refund, I just need a copy of my receipt."},"expected":{"label":"other"}}

The function returns, for each case, an object such as {"answer": "Use password reset, ...", "label": "account"}.

Choosing the evaluators

CriterionEvaluatorNeeds expectedWhat it measures
exact_labelexact_match on labelyestask success: the label is the right one
format_validjson_schema over the whole outputnoformat: the object has exactly the two string fields
pii_freeregex on answer, pass_if: no_matchnoa safety property of the text

A format check passes a well-formed wrong answer, so it can never stand in for task success. A task check needs a reference for every case; where a case has none, exact_match cannot score it. Deterministic evaluators need no validation against people: running one twice gives the same verdict.

The policy

release.yaml is what the run is decided against:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
warn_on: []
rules:
  - id: exact-label-floor
    metric: exact_label
    min: 0.70
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: pii-free
    metric: pii_free
    kind: observed_count
    max_failures: 0

exact-label-floor says the label must be right at least 70% of the time, and passes only when the whole 95% interval is at or above 0.70. The two observed_count rules allow no failure at all on the cases you ran; they describe these cases, not every question users will ask.

Run it

oloproof run

Real output, trimmed:

Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ exact-label-floor │ exact_label  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ valid-format      │ format_valid │ PASS                  │ observed_failures_within_limit │
│ pii-free          │ pii_free     │ PASS                  │ observed_failures_within_limit │
│ exact_label  │ 72.2%    │ [46.5%, 90.4%]  │ 13 / 18 observed · 0 missing · 0 excluded │
│ format_valid │ 100.0%   │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
│ pii_free     │ 100.0%   │ [81.4%, 100.0%] │ 18 / 18 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/18 miss; judgment 0 hit/54 miss

How to read it:

  • 13 of 18 labels are right, 72.2%. That is above 0.70, but the interval reaches down to 46.5%: 18 cases cannot show the true rate is at least 0.70. So the rule is INSUFFICIENT_EVIDENCE, not PASS and not FAIL.
  • Every output has the right shape and none contains an SSN-like number, so both format rules pass.
  • block_on lists INSUFFICIENT_EVIDENCE, so the gate blocks and the command exits 3. Exit 0 would mean nothing the policy blocks on; Gating CI lists every code.

Run it again and the cache line reads execution 18 hit/0 miss: nothing changed, so the function is not called.

Inspect the failures

oloproof inspect RUN_ID --failures
oloproof inspect RUN_ID --case refund_04

RUN_ID is the id on the first line of the run's output.

5 of 18 cases failed, errored or did not finish

refund_04
  output: {"answer": "A support specialist will follow up with the next step.", "label": "other"}
  exact_label: failed
...
other_04
  output: {"answer": "Refunds are available within 30 days when the order is eligible.", "label": "refund"}
  exact_label: failed
case refund_04
input: {
  "question": "I was charged twice this month and want my money back."
}
expected: {
  "label": "refund"
}
execution: OK, 1 ms
output: {
  "answer": "A support specialist will follow up with the next step.",
  "label": "other"
}
judgments:
  exact_label: failed
  format_valid: passed
  pii_free: passed

The pattern is plain once you read the inputs: "money back", "reverse the payment", "sign in" and "two-factor" are not in the keyword lists, and other_04 says "I don't want a refund", which the word "refund" matches anyway. Note that refund_04 passes both format checks while being wrong: that is the gap between checking format and measuring success.

Two next actions are meaningful here. Fix the misses (below), or add cases: with more cases at the same accuracy the interval narrows, and oloproof plan RUN_ID --run estimates how many.

Make a real change

Copy ../support-change/app.py over app.py. It adds the missed phrasings:

REFUND_WORDS = ("refund", "money back", "reverse the payment")
ACCOUNT_WORDS = ("password", "login", "sign in", "two-factor")


@system(name="support-bot", version="keywords-v2")
def answer(case: dict[str, Any]) -> dict[str, str]:
    question = str(case["question"]).lower()
    if any(word in question for word in REFUND_WORDS):
        ...

and set version: keywords-v2 under system in oloproof.yaml, so the run is recorded as the new version. Then:

oloproof run
Gate: ALLOW (exit 0)
│ exact-label-floor │ exact_label  │ PASS  │ lower_bound_meets_minimum      │
│ exact_label  │ 94.4%    │ [72.7%, 99.9%]  │ 17 / 18 observed · 0 missing · 0 excluded │

17 of 18 are right and the interval's lower bound, 72.7%, clears 0.70, so the rule passes and the command exits 0. other_04 still fails: the fix did not touch negation.

Compare the candidate with the baseline

The run rule asks whether the candidate meets your floor. A comparison asks how it differs from the baseline, case by case. Copy ../support-change/compare.yaml into the project; it holds one comparison rule:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: label-no-regression
    kind: non_inferiority
    metric: exact_label
    margin: 0.10
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
Comparison sha256:de76... of run_01M4...TJAD against run_01M4...ECVEF · 18 paired cases
exact_label: +22.2 points [-12.9, +57.0] · 18 paired · 0 missing · 0 excluded
format_valid: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
pii_free: +0.0 points [-25.8, +25.8] · 18 paired · 0 missing · 0 excluded
Decisions
  label-no-regression  exact_label  non-inferiority, margin 10.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    about 3 more paired cases would decide it, if the difference holds (21 in total at 22% discordance)
Gate: BLOCK (exit 3)

The candidate fixed four cases and broke none, an estimated gain of 22 points. But only four cases changed, and 18 paired cases leave an interval from 12.9 points worse to 57 points better, which crosses the 10-point margin. The comparison cannot yet rule out that the candidate is worse by more than you accept, so it is INSUFFICIENT_EVIDENCE and exits 3. The line beneath it is the sizing estimate. Without --policy, compare prints the differences, says the project's release.yaml declares no comparison rule, and exits 0 because nothing was decided.

Comparing a candidate to a baseline explains the margin and the other rule kinds.

Troubleshooting

SymptomCause and fix
ModuleNotFoundError for appRun from the directory holding app.py, or give callable a module path importable from there.
A rule names a metric no evaluator producesThe rule's metric must equal an evaluator's criterion; the error lists the metrics there are.
You edited the classifier and the run reused every outputThe cache follows the callable's source; a helper file it reads must be listed under system.code_paths.
exact_label reports cases as missingThose executions raised or timed out; oloproof inspect RUN_ID --failures shows each error.
The run exits 3 with a high estimateThe interval, not the estimate, decides. Add cases or accept a lower floor, decided before the run.

Limitations

  • The SDK and YAML report pass rates per evaluator. There is no confusion matrix or per-class precision and recall for a classifier like this one; a predictive: block does those for a scored model (Predictive models).
  • observed_count rules describe the cases you ran; they make no claim about unseen inputs.
  • A comparison over 18 cases resolves only large differences. Fifty or more real cases are a more useful floor.
  • oloproof.yaml cannot name a custom @evaluator; that needs the SDK (The SDK).