Skip to content

Guides

Tutorial: a classifier and a regressor

A runnable walkthrough for three predictive models evaluated from the command line: a binary churn classifier, a three-class ticket router evaluated one class at a time, and a delivery-time regressor scored by absolute error within a declared range. Each runs locally with no provider, no key and no network, and each ends with a candidate change measured against the baseline.

This page is the hands-on companion to Classifiers and regressors, which explains every field; read that page for the reference and this one to do it once end to end. Terms such as interval, decision state and release action are defined in Concepts.

What Oloproof measures for a predictive model, and what it does not

ModelWhat you getWhat you do not get
Binary classifieraccuracy, recall, precision, Brier score, log loss, ROC-AUC and average precision, plus confusion counts, a calibration table and a threshold sweepa recommended threshold
Multiclass classifierone recall and one precision per class, each a binary metric whose positive class is that class (one against the rest)a macro or micro average metric
Regressormean absolute error, bounded by a target_range you declaresquared error, R squared, or any error without a declared range

The model itself stays yours. Oloproof calls a Python function you point it at, reads the prediction it returns, and never sees features, weights or internals.

Prerequisites

  • Python 3.11 or later and pip install oloproof in a virtual environment, as in the quickstart.
  • The example files, which ship with the package. Copy them into a new directory so the runs' stores land there:
oloproof init --example predictive ~/oloproof-predictive
cd ~/oloproof-predictive

Every command below runs from one of its three subdirectories. Each run writes its evidence to a .oloproof/ directory beside the oloproof.yaml it read.

predictive/
  binary/        app.py  oloproof.yaml  candidate.yaml  release.yaml  comparison.yaml  data/accounts.jsonl
  multiclass/    app.py  oloproof.yaml  candidate.yaml  release.yaml  data/tickets.jsonl
  regression/    app.py  oloproof.yaml  candidate.yaml  release.yaml  data/orders.jsonl

Part 1: a binary classifier

The adapter

binary/app.py holds the model and the adapter in one file. The model is the same deterministic churn model as the churn_model example (oloproof init --example churn_model); the adapter is the decorated function Oloproof calls:

from oloproof import system


@system(name="churn-model", version="baseline")
def run(account):
    score = churn_score(account)
    return {"label": score >= 0.5, "score": round(score, 4)}


@system(name="churn-model", version="candidate-cutoff-0.4")
def run_candidate(account):
    score = churn_score(account)
    return {"label": score >= 0.4, "score": round(score, 4)}

To evaluate your own model, keep the shape and replace churn_score with a call to it, for example model.predict_proba([features(account)])[0][1] for a scikit-learn model you load once at import. Change version whenever the model or its cut-off changes: the version is part of the cache key, so a model retrained under the same version would reuse the old predictions.

Input and output shapes

One line of data/accounts.jsonl is one case:

{"expected": {"label": false}, "id": "account_000", "input": {"recent_upgrade": true, "support_contacts": 0, "tenure_months": 0}, "metadata": {"plan": "enterprise"}}
PartShapeWho reads it
inputthe object your function receivesyour adapter
expected.labeltrue or false, the truththe evaluators
metadata.planany JSONslices only
returned labeltrue or false, the predictionthe evaluators
returned scorea number in [0, 1], the probability of the positive classBrier, log loss, ranking, calibration, sweep

The evaluators, and why these

binary/oloproof.yaml:

version: 1
project: churn-tutorial
dataset: data/accounts.jsonl
system:
  name: churn-model
  version: baseline
  callable: app:run
evaluators:
  - {type: predictive_correct, criterion: accuracy}
  - {type: predictive_recall, criterion: recall}
  - {type: predictive_precision, criterion: precision}
  - {type: predictive_brier, criterion: brier}
  - {type: predictive_log_loss, criterion: log_loss, clip: 0.02}
  - {type: predictive_ranking, criterion: rank}
metrics:
  - {id: roc_auc, type: ranking, criterion: rank, statistic: roc_auc}
  - {id: pr_auc, type: ranking, criterion: rank, statistic: average_precision}
predictive:
  label_field: label
  score_field: score
  expected_field: label
  positive: true
  calibration_bins: 10
  thresholds: [0.3, 0.4, 0.5, 0.6, 0.7]
slices: [metadata.plan, "confidence:0.5"]
min_slice_support: 20
  • Accuracy, recall and precision are three rates over three different sets of rows: every account, the accounts that churned, and the accounts the model flagged. A churn model needs all three because about a third of accounts churn, so a model that predicts nobody churns is 68% accurate and finds no one.
  • Brier and log loss score the probability behind the label. Log loss needs clip, because one confident mistake is otherwise infinite.
  • predictive_ranking plus two metrics: entries give ROC-AUC and average precision. They answer how well the model orders accounts, separately from where the cut-off sits.
  • The predictive: block says where the label, score and truth live, and produces the confusion counts, the calibration table and the threshold sweep beside the metrics.

The policy

binary/release.yaml puts a floor on each rate:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - {id: accuracy-floor, metric: accuracy, min: 0.75}
  - {id: recall-floor, metric: recall, min: 0.60}
  - {id: precision-floor, metric: precision, min: 0.60}

A rule passes when the interval's lower bound clears the floor, fails when the upper bound is below it, and otherwise reads INSUFFICIENT_EVIDENCE.

Run it

cd binary
oloproof run

Trimmed:

Run run_... [DECIDED/COMPLETE]
Gate: ALLOW (exit 0)
│ accuracy-floor  │ accuracy  │ PASS  │ lower_bound_meets_minimum │
│ recall-floor    │ recall    │ PASS  │ lower_bound_meets_minimum │
│ precision-floor │ precision │ PASS  │ lower_bound_meets_minimum │

│ accuracy  │ 88.5%    │ [83.2%, 92.6%]  │ 177 / 200 observed · 0 missing · 0 excluded                                │
│ recall    │ 81.2%    │ [69.5%, 90.0%]  │ 52 / 64 observed · 0 missing · 136 excluded                                │
│ precision │ 82.5%    │ [70.9%, 91.0%]  │ 52 / 63 observed · 0 missing · 137 excluded                                │
│ brier     │ 0.120    │ [0.094, 0.154]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ log_loss  │ 0.389    │ [0.323, 0.499]  │ mean of 200 observed · 0 missing · 0 excluded                              │
│ roc_auc   │ 92.3%    │ [69.3%, 100.0%] │ roc_auc over 64 positive · 136 negative · 0 missing · 0 excluded           │
│ pr_auc    │ 86.5%    │                 │ average_precision over 64 positive · 136 negative · 0 missing · 0 excluded │

How to read it:

  • Gate: ALLOW (exit 0) means no rule reached a state the policy blocks on. It is not a claim that the model is good beyond the three floors you wrote.
  • The excluded counts are the denominators at work: recall is measured over the 64 accounts that churned, so the 136 that did not are excluded from it, not counted as failures.
  • pr_auc has no interval. At 200 rows the engine withholds a bound it cannot support; a rule on it would read INSUFFICIENT_EVIDENCE with interval_unavailable.

Below the metrics, the same output prints the confusion counts, the calibration table and the threshold sweep:

│ actually positive │ 52                 │ 12                 │
│ actually negative │ 11                 │ 125                │

│ 0.2-0.3 │ 26.7%   │ 0.0%     │ 34 rows │
│ 0.5-0.6 │ 53.4%   │ 81.0%    │ 21 rows │

│ 0.3     │ 57.4%     │ 96.9%  │ 62/108 predicted positive · 62/64 actual positive │
│ 0.4     │ 62.6%     │ 89.1%  │ 57/91 predicted positive · 57/64 actual positive  │
│ 0.5     │ 82.5%     │ 81.2%  │ 52/63 predicted positive · 52/64 actual positive  │
│ 0.6     │ 83.3%     │ 54.7%  │ 35/42 predicted positive · 35/64 actual positive  │
│ 0.7     │ 100.0%    │ 37.5%  │ 24/24 predicted positive · 24/64 actual positive  │

The counts are not rates and no rule may name them. The calibration rows say the model's probabilities are off: in the 0.2 to 0.3 band it claims about one churner in four and none of 34 churned. The sweep is titled Thresholds (exploratory; recommends nothing): it shows what each cut-off would have measured and leaves the choice to you, because only you know what a missed churner costs against a wasted retention call.

Inspect the failures

oloproof inspect RUN_ID --failures --limit 5
23 of 200 cases failed, errored or did not finish

account_032
  output: {"label": false, "score": 0.07}
  accuracy: failed
  recall: failed

account_037
  output: {"label": false, "score": 0.37}
  accuracy: failed
  recall: failed
...

RUN_ID is the id on the first line of the run's output. oloproof inspect RUN_ID --case account_037 shows one case's input, expected value, output and every judgment. Several missed churners sit just under the 0.5 cut-off (0.37, 0.43), which is what the sweep's 0.4 row already suggested.

A meaningful next action follows from what you see, not from the gate: here the misses cluster below the cut-off, so the candidate tries a lower one. Had they been confident misses (0.07), the next action would be the model's features, and no cut-off would help.

A candidate change, and the comparison

binary/candidate.yaml is oloproof.yaml with two lines changed:

system:
  name: churn-model
  version: candidate-cutoff-0.4
  callable: app:run_candidate
oloproof run --config candidate.yaml
Gate: BLOCK (exit 3)
│ accuracy-floor  │ accuracy  │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ recall-floor    │ recall    │ PASS                  │ lower_bound_meets_minimum   │
│ precision-floor │ precision │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │

│ accuracy  │ 79.5%    │ [73.2%, 84.9%]  │ 159 / 200 observed · 0 missing · 0 excluded                                │
│ recall    │ 89.1%    │ [78.7%, 95.5%]  │ 57 / 64 observed · 0 missing · 136 excluded                                │
│ precision │ 62.6%    │ [51.8%, 72.6%]  │ 57 / 91 observed · 0 missing · 109 excluded                                │

Exactly the sweep's 0.4 row: recall up, precision down. Exit 3 is INSUFFICIENT_EVIDENCE, not FAIL: the floors are inside the intervals, so 200 accounts cannot say which side the candidate is on. The scores did not change, so Brier, log loss and ROC-AUC are identical.

Now compare the two runs case by case. binary/comparison.yaml holds comparison rules:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - {id: recall-no-worse, kind: non_inferiority, metric: recall, margin: 0.05}
  - {id: accuracy-no-worse, kind: non_inferiority, metric: accuracy, margin: 0.05}
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy comparison.yaml
Comparison sha256:... of run_... against run_... · 200 paired cases
accuracy: -9.0 points [-17.9, -1.4] · 200 paired · 0 missing · 0 excluded
recall: +7.8 points [-2.8, +21.9] · 64 paired · 0 missing · 136 excluded
  excluded 136: not_a_positive_case
precision: +0.0 points [-8.3, +8.3] · 63 paired · 0 missing · 137 excluded
  excluded 137: not_predicted_positive
...
Decisions
  recall-no-worse  recall  non-inferiority, margin 5.0 points  PASS  lower_bound_above_margin
  accuracy-no-worse  accuracy  non-inferiority, margin 5.0 points  INSUFFICIENT_EVIDENCE  interval_overlaps_margin
    no sample size would make this PASS: the difference itself (-9.0 points) is outside the margin, so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

Read the precision line carefully. The candidate's own precision fell from 82.5% to 62.6%, yet the paired difference is +0.0 over 63 pairs. A comparison pairs rows: a precision difference is measured only over accounts both models flagged, and on those 63 both were right. The 28 extra accounts the candidate flagged, 23 of them false alarms, are outside that set. That is why comparison.yaml guards accuracy rather than precision for a cut-off change; Comparison rules has the rule kinds.

The decision is the useful part: recall is shown no worse, and accuracy cannot be shown within five points; the advice line says more data would push it toward FAIL. Whether trading nine points of accuracy for eight of recall is worth it is a business decision the gate has made visible.

Part 2: a multiclass classifier, one class at a time

multiclass/app.py routes a support ticket to one of three queues, and has one deliberate flaw: any ticket sent from the mobile app goes to technical.

@system(name="ticket-router", version="baseline")
def route(ticket):
    if ticket["channel"] == "app":
        return {"queue": "technical"}
    return {"queue": classify(str(ticket["subject"]))}

A case:

{"expected": {"queue": "billing"}, "id": "ticket_000", "input": {"channel": "email", "subject": "update the card on file"}, "metadata": {"channel": "email"}}

There is no multiclass metric to switch on. Each class gets its own binary block: recall for billing is the binary recall whose positive class is billing. multiclass/oloproof.yaml:

evaluators:
  - {type: predictive_correct, criterion: accuracy, field: queue, expected_field: queue}
  - {type: predictive_recall, criterion: recall_billing, field: queue, expected_field: queue, positive: billing}
  - {type: predictive_precision, criterion: precision_billing, field: queue, expected_field: queue, positive: billing}
  - {type: predictive_recall, criterion: recall_technical, field: queue, expected_field: queue, positive: technical}
  - {type: predictive_precision, criterion: precision_technical, field: queue, expected_field: queue, positive: technical}
  - {type: predictive_recall, criterion: recall_account, field: queue, expected_field: queue, positive: account}
  - {type: predictive_precision, criterion: precision_account, field: queue, expected_field: queue, positive: account}
slices: [metadata.channel]

Oloproof does not compute a macro or micro average over these. If you need one, it is a number you derive yourself from the per-class counts, and no rule can gate on it. The policy puts a floor on the classes that matter, because a router can be accurate overall and lose one queue:

rules:
  - {id: billing-recall-floor, metric: recall_billing, min: 0.80}
  - {id: account-recall-floor, metric: recall_account, min: 0.80}
  - {id: technical-precision-floor, metric: precision_technical, min: 0.80}
cd ../multiclass
oloproof run
Gate: BLOCK (exit 1)
│ billing-recall-floor      │ recall_billing      │ FAIL                  │ upper_bound_below_minimum   │
│ account-recall-floor      │ recall_account      │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ technical-precision-floor │ precision_technical │ FAIL                  │ upper_bound_below_minimum   │

│ accuracy            │ 81.3%    │ [74.1%, 87.3%]  │ 122 / 150 observed · 0 missing · 0 excluded │
│ recall_billing      │ 66.0%    │ [51.2%, 78.8%]  │ 33 / 50 observed · 0 missing · 100 excluded │
│ precision_billing   │ 100.0%   │ [89.4%, 100.0%] │ 33 / 33 observed · 0 missing · 117 excluded │
│ recall_technical    │ 100.0%   │ [92.8%, 100.0%] │ 50 / 50 observed · 0 missing · 100 excluded │
│ precision_technical │ 64.1%    │ [52.4%, 74.7%]  │ 50 / 78 observed · 0 missing · 72 excluded  │
│ recall_account      │ 78.0%    │ [64.0%, 88.5%]  │ 39 / 50 observed · 0 missing · 100 excluded │
│ precision_account   │ 100.0%   │ [90.9%, 100.0%] │ 39 / 39 observed · 0 missing · 111 excluded │

Exit 1 means at least one rule is FAIL. Each class has its own denominator: 78 tickets were called technical, so that is precision's denominator for technical. The slice table points at the cause; slices are exploratory and never gate, but a marked slice is a lead worth reading:

│ metadata.channel=app   │ accuracy            │ 42.9%    │ [28.8%, 57.8%]  │ 21 / 49 observed · ... │ marked (p 0.0001)    │
oloproof inspect RUN_ID --failures --limit 2
28 of 150 cases failed, errored or did not finish

ticket_011
  output: {"queue": "technical"}
  accuracy: failed
  precision_technical: failed
  recall_account: failed
...

The candidate (route_candidate, run with oloproof run --config candidate.yaml) drops the channel shortcut. On this synthetic data it routes every ticket correctly and the gate allows:

Gate: ALLOW (exit 0)
│ billing-recall-floor      │ recall_billing      │ PASS  │ lower_bound_meets_minimum │
│ account-recall-floor      │ recall_account      │ PASS  │ lower_bound_meets_minimum │
│ technical-precision-floor │ precision_technical │ PASS  │ lower_bound_meets_minimum │

oloproof compare works here exactly as in Part 1, one difference per class metric.

Part 3: a regressor, scored by absolute error

regression/app.py estimates delivery days and ignores whether the item is in stock:

@system(name="delivery-estimator", version="baseline")
def estimate(order):
    return {"days": round(base_days(order), 1)}

A case, with the truth as a number:

{"expected": {"days": 11}, "id": "order_000", "input": {"distance_km": 468, "express": false, "in_stock": false}, "metadata": {"in_stock": false}}

The only regression evaluator is absolute error, and it needs the range every target lies in. An absolute error is a bounded mean whose interval holds only within that range, so it is declared, never defaulted. Deliveries here take 0 to 20 days:

evaluators:
  - type: predictive_absolute_error
    criterion: days_error
    field: days
    expected_field: days
    target_range: [0, 20]
slices: [metadata.in_stock]

The policy is a budget, a max: rule: it passes when the interval's upper bound is at or below it.

rules:
  - {id: error-budget, metric: days_error, max: 2.0}
cd ../regression
oloproof run
Gate: BLOCK (exit 1)
│ error-budget │ days_error │ FAIL  │ lower_bound_above_maximum │

│ days_error │ 2.66     │ [2.01, 3.57] │ mean of 120 observed · 0 missing · 0 excluded │

│ metadata.in_stock=false │ days_error │ 5.18     │ [4.60, 6.65] │ mean of 53 observed · ... │
│ metadata.in_stock=true  │ days_error │ 0.67     │ [0.51, 2.19] │ mean of 67 observed · ... │

FAIL because even the interval's lower bound is above the two-day budget. A score evaluator has no pass or fail per case, so oloproof inspect RUN_ID --failures lists none; read a case instead:

oloproof inspect RUN_ID --case order_000
output: {
  "days": 6.1
}
judgments:
  days_error: score 4.9

The slice says where to look: out-of-stock orders are off by five days. The candidate (estimate_candidate) adds five days for an item out of stock:

oloproof run --config candidate.yaml
Gate: ALLOW (exit 0)
│ error-budget │ days_error │ PASS  │ upper_bound_meets_maximum │
│ days_error │ 0.70     │ [0.56, 1.57] │ mean of 120 observed · 0 missing · 0 excluded │

Troubleshooting

SymptomCause and fix
Configuration error: evaluator 'recall' counts 'churned' as the positive class, and no case's 'label' is 'churned'positive: names a value no case has. Use the value exactly as it appears under expected, including true versus "true".
release rule ... refers to unknown metric 'false_positives'Confusion counts are not metrics. Gate on recall or precision.
A regression evaluator is refused before the runtarget_range is missing or empty. Declare the range the targets can actually take; a wider range gives a wider interval.
A ranking rule reads INSUFFICIENT_EVIDENCE (interval_unavailable)The suite is too small for that statistic's interval, typically pr_auc. Gate on roc_auc, or add cases.
The candidate reports the baseline's numbersBoth runs share a version, so the cached predictions were reused. Give every change its own version.
A case is missing rather than wrongYour function raised, or the field the evaluator reads is absent or not a number. oloproof inspect RUN_ID --case ID shows the error.

Limitations

  • No recommended threshold. The sweep reports what each declared cut-off would have measured.
  • No macro or micro average metric for multiclass, and no multi-label support beyond one block per label.
  • Regression is absolute error within a declared target_range only: no squared error, R squared or unbounded error.
  • Calibration is shown as a table, not gated and not corrected.
  • A paired precision difference covers only rows both models flagged, as Part 1 shows.
  • The examples are deterministic and synthetic. Their decisions show the mechanics, not how a real model behaves on real data.

Where to go next