Guides
Tutorial: text generation with a rubric judge
Evaluate a function that writes free text, here a ticket summariser, with format checks and a rubric judge; measure that judge against a person's labels before it may decide anything; then compare a real change. The judge runs on this machine with no model and no network, and an optional step swaps in a real model.
What you will build
A summariser that turns a support ticket into one or two sentences. "Good" is a judgment, not a string match, so task success is decided by an LLM judge with a rubric: does the summary state the facts an agent needs? Two deterministic evaluators check format, which needs no reference. Terms such as case, run, metric, judge and gate are defined in Concepts.
The same shape fits extraction or any other generation: a function returns text in a dictionary, the reference says what a good answer must contain, and a rubric says how to decide.
Prerequisites
- Python 3.11 or later, and Oloproof in a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof- The example project, which ships with the package. Copy it into a new directory and work there:
oloproof init --example generation ticket-summaries
cd ticket-summaries- Port 8799 free for the stand-in judge (change it in both places if not).
Every step until "Optional: a real model as the judge" is offline and deterministic: no API key, no provider account, no cost.
The files
ticket-summaries/
app.py the summariser under test (baseline)
app_v2.py the candidate change
judge_server.py a stand-in judge speaking the OpenAI API on 127.0.0.1
rubrics/covers_facts.md the judge's rubric
oloproof.yaml the suite
release.yaml rules for a run
compare.yaml a rule for a comparison
data/tickets.jsonl 20 cases
labels/reviewer_verdicts.csv one person's verdicts on the baseline's summaries
fill_labels.py copies those verdicts into a labelling sheetRun every command from ticket-summaries/.
The stand-in judge, and what it is not
A rubric judge is an evaluator that sends a prompt (the rubric, the case's input, its expected and the output) to a model and reads back {"pass": true|false, "rationale": "..."}. Oloproof talks to any server speaking the OpenAI chat API, and a server on localhost needs no key.
judge_server.py is such a server, but it is not a model. It passes a summary only when it contains every phrase under must_mention in the case's expected, ignoring case. That is a fixed rule, so the tutorial gives the same numbers on every machine. It cannot notice an invented fact, which a real model judge is asked to. Start it in a second terminal and leave it running:
python judge_server.py --port 8799stand-in judge on http://127.0.0.1:8799/v1The application and its adapter
# app.py
@system(name="ticket-summariser", version="first-sentence")
def summarise(case: dict[str, Any]) -> dict[str, str]:
return {"summary": sentences(str(case["ticket"]))[0]}The adapter for a Python application is the function: it receives the case's input and returns a dictionary. For your own generator, call your model or chain inside it and return the text under a key. Oloproof calls it once per case and caches the output on the function's source and declared version; it does not manage your model client, prompts or state. List files the function reads, such as a prompt template, under system.code_paths.
The dataset
{"id":"t01","input":{"ticket":"Hello. Order 1042 arrived with a cracked screen. I would like a replacement, not a refund."},"expected":{"must_mention":["1042","cracked","replacement"]}}
{"id":"t06","input":{"ticket":"Please cancel my subscription at the end of this month. I am moving abroad."},"expected":{"must_mention":["cancel","end of this month"]}}input is what the function receives. expected is the reference the judge reads: here a list of facts the summary must carry, not a full reference summary, because many different summaries are correct. The output for t01 is {"summary": "Hello."}.
Choosing the evaluators
version: 1
project: ticket-summaries
dataset: data/tickets.jsonl
system:
name: ticket-summariser
version: first-sentence
callable: app:summarise
timeout_s: 30
evaluators:
- type: json_schema
criterion: format_valid
field: null
schema:
type: object
required: [summary]
properties:
summary: {type: string, minLength: 1}
additionalProperties: false
- type: regex
criterion: short_enough
field: summary
pattern: '^.{1,160}$'
pass_if: match
- type: rubric_judge
criterion: covers_facts
provider: openai_compatible
model: stand-in-judge
base_url: http://127.0.0.1:8799/v1
rubric_file: rubrics/covers_facts.md| Criterion | Evaluator | Needs expected | Measures |
|---|---|---|---|
| format_valid | json_schema | no | format: one non-empty string field |
| short_enough | regex | no | format: at most 160 characters |
| covers_facts | rubric_judge | yes | task success, as the rubric defines it |
Hello. passes both format checks. Only the judge says it is a useless summary. A judge can also run without a reference: a rubric such as "PASS if the summary contains no greeting" reads only the input and output, and a case with no expected is still judged. What it cannot do then is check facts against an answer you trust.
The rubric:
PASS when the summary states every fact listed under must_mention in the expected answer, in
words a support agent would recognise, and adds nothing the ticket does not say.
FAIL when any listed fact is missing, changed or contradicted.The policy
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
- id: valid-format
metric: format_valid
kind: observed_count
max_failures: 0
- id: short-enough
metric: short_enough
kind: observed_count
max_failures: 0
- id: covers-facts-floor
metric: covers_facts
min: 0.60require_validated_evaluators: true is the engine's default, written out here because it is the point of this tutorial: a judge nobody has compared with people may not decide a rule.
Run it
oloproof runRun run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format │ format_valid │ PASS │ observed_failures_within_limit │
│ short-enough │ short_enough │ PASS │ observed_failures_within_limit │
│ covers-facts-floor │ covers_facts │ INSUFFICIENT_EVIDENCE │ evaluator_not_validated │
covers-facts-floor: the judge (or model or custom evaluator) behind this rule has not been measured against
people yet, so it may not decide.
Label a sample: oloproof review run_01M4... --criterion covers_facts --by YOU --sample 20
Then measure it: oloproof evaluators validate EVALUATOR_ID --by YOU (ids: oloproof evaluators list)
│ format_valid │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ short_enough │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ covers_facts │ 45.0% │ [23.0%, 68.5%] │ 9 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 missThe format rules pass. The judge passed 9 of 20 summaries, but the rule is INSUFFICIENT_EVIDENCE with the reason evaluator_not_validated, and the gate blocks with exit 3. The rule did not decide on the 45%: a judge's error rate is unknown until it is measured, so an interval built on its verdicts would carry an unstated error. The engine reports this as INSUFFICIENT_EVIDENCE, not MANUAL_REVIEW or FAIL: the evidence to decide is missing, and the output prints the two commands that supply it.
Inspect the failures
oloproof inspect RUN_ID --failures11 of 20 cases failed, errored or did not finish
t01
output: {"summary": "Hello."}
covers_facts: failed
judge text, not verified: missing: 1042, cracked, replacement
t02
output: {"summary": "I was charged twice for order 2210."}
covers_facts: failed
judge text, not verified: missing: 49
...The judge's rationale is shown as "judge text, not verified": it is the model's explanation, not evidence. The pattern is clear anyway: the first sentence is often a greeting.
Measure the judge against a person
Validation compares the judge's verdicts with a person's on the same answers. Draw a random sample of the run's cases into a sheet. The judge's verdicts are left out of it, so the labeller is not anchored on them:
oloproof labels export RUN_ID --criterion covers_facts --sample 20 --local --out sample.csvWrote 20 cases to sample.csv, drawn at random with seed 2701013296, without the judge's verdict.
This is a local sample, good-faith only, because it was drawn on this machine.
Fill in `passed` (pass or fail) and `labelled_by` on each row you judge, then run `oloproof labels import sample.csv`.--local draws on this machine without asking a hosted workspace; the engine chooses the seed. With 20 cases, a sample of 20 is all of them. In practice, a person reads each row's ticket and summary and fills in passed. For this tutorial, labels/reviewer_verdicts.csv holds verdicts a reviewer gave on the baseline's summaries, and fill_labels.py copies them into the sheet:
python fill_labels.py sample.csv
oloproof labels import sample.csvfilled 20 rows of sample.csv
Recorded 20 labels from sample.csv (20 measurement).The reviewer disagreed with the judge once: on t02 ("I was charged twice for order 2210.") they judged the missing amount immaterial and passed it. Labels name the exact answer they judged, so these verdicts apply to the baseline run only.
Find the judge's version id and validate it:
oloproof evaluators list
oloproof evaluators validate EVALUATOR_ID --by alicecovers_facts LLM_JUDGE UNVALIDATED (declared) sha256:a662...
covers_facts: sha256:a662... is now VALIDATED
agreement 95.0% [75.1%, 99.9%] · 19 of 20 labelled cases agreed · 0 labelled but not judged · kappa 0.900
bias -5.0 points [-32.4, +20.7] · the judge's pass rate minus the people's · 20 cases · 0 labelled but not judged
passes what people pass 90.0% [55.4%, 99.8%] · the judge passed 9 of 10 cases people passed · 0 labelled but not judged
fails what people fail 100.0% [69.1%, 100.0%] · the judge failed 10 of 10 cases people failed · 0 labelled but not judgedRead the intervals, not the 95%: 20 labels show agreement of at least 75.1%. A policy can demand more with minimum_evaluator_agreement, which compares that lower bound, and validate refuses a judge below it. The Judges guide covers the bar, bias, probes and oloproof review for labelling in the terminal.
Now re-decide the stored run without calling the summariser or the judge:
oloproof gate RUN_ID --policy release.yamlvalid-format: PASS (observed_failures_within_limit)
short-enough: PASS (observed_failures_within_limit)
covers-facts-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)
no sample size would make this PASS: the observed rate (0.500) is itself below the threshold (0.600), so more cases would move it toward FAIL
Gate: BLOCK (exit 3)The judge may decide now, and the decision is about the summariser: the rate it quotes, 0.500, is not the judge's 45%. Because this run has a blind, random sample of measurement labels, the gate reads the judge corrected by those labels ("Judge-corrected gates" in the Judges guide). The correction is PPI, prediction-powered inference: it uses the labelled sample to measure how far the judge's rate sits from the people's, and moves the estimate and widens the interval by that much. That is also what the export's notes about PPI refer to. Either way, the baseline does not meet the floor, and more cases would not change that.
Make a real change
app_v2.py skips short pleasantries and keeps the next two sentences. Copy it over app.py, set version: skip-pleasantries under system in oloproof.yaml, keep the judge running, and:
oloproof runGate: ALLOW (exit 0)
│ covers-facts-floor │ covers_facts │ PASS │ lower_bound_meets_minimum │
│ covers_facts │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 6 hit/54 missThe judge is the same validated version, so its rule decides directly. Six judgments came from the cache, on summaries both versions wrote identically. Nobody labelled these new summaries; the judge's validation is what lets its verdicts stand.
Compare the candidate with the baseline
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
- id: covers-more-facts
kind: superiority
metric: covers_factsoloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlformat_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
short_enough: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
covers_facts: +55.0 points [+13.0, +84.4] · 20 paired · 0 missing · 0 excluded
Decisions
covers-more-facts covers_facts superiority PASS difference_above_zero
Gate: ALLOW (exit 0)A comparison does not apply the PPI correction: it compares the judge's own verdicts on the two runs, which is why the gain starts from the judge's 45% and not from the corrected 0.500 above. Eleven summaries improved and none got worse; the interval for the gain is entirely above zero, so the superiority rule passes and the command exits 0. Format is guarded by the run rules, which allow no failure, rather than by a comparison: over 20 cases a comparison of two perfect format scores could only say the difference is within 23.6 points.
Optional: a real model as the judge
This step leaves the offline path. It needs a model server, and with a cloud provider a key and money.
- Local, no key and no cost: Ollama, LM Studio or llama.cpp on localhost. Pull a chat model (for Ollama, ollama pull llama3.1).
- Cloud: provider: anthropic or openai with api_key_env naming the variable that holds your key, or openai_compatible with base_url and api_key_env. Each case is one judge call (two when the first reply is not valid JSON), billed at your provider's rates, and Oloproof never calls a judge again for an answer it has already judged.
Write the draft judge in a file of its own, as it would appear under evaluators::
# live_judge.yaml
type: rubric_judge
criterion: covers_facts
provider: openai_compatible
model: llama3.1
base_url: http://localhost:11434/v1
rubric_file: rubrics/covers_facts.mdand try it against the answers your reviewer already labelled, without validating or adopting it:
oloproof evaluators try live_judge.yamlLocal servers answer one request at a time by default; add concurrency: {system: 2, judge: 2} to oloproof.yaml so queued calls do not time out. A run of this step with a small local model (qwen2.5vl) on a laptop printed:
covers_facts: draft sha256:b88a... on 20 labelled cases · 20 judged now, 0 from cache, 11 errored
agreement 88.9% [19.1%, 99.9%] · 8 of 9 labelled cases agreed · 11 labelled but not judged · kappa 0.769Eleven calls timed out, and the agreement interval counts each one both ways, so it reaches down to 19.1%: a judge that does not answer is not measured. A larger model, a longer timeout, or fewer concurrent calls is the fix. To adopt the model, put it in oloproof.yaml in place of the stand-in. That is a new evaluator version: its configuration (model, endpoint, rubric) is its identity, so the stand-in's validation does not carry over. Run the baseline again with it and validate it against the labels, as above.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| covers_facts all missing, no_observations | The judge server is not running or not on base_url. Every judge call errored; oloproof inspect RUN_ID --failures shows why. |
| evaluator_not_validated after you validated | You changed the judge (model, endpoint, port, rubric) and made a new version. Validate that one. |
| labels import refuses the file and names a row | The row names a case or execution the run does not hold; export again from the run you label. |
| labels export says a workspace could not be reached | You are logged in to one, so it asked it to draw. --local draws here instead. |
| A cloud judge fails before any call | Its key is not in the variable api_key_env names. |
Limitations
- The stand-in judge is a phrase match. It demonstrates the workflow, not judging quality.
- There are no BLEU, ROUGE or embedding-similarity evaluators. In the SDK, write one with @evaluator; oloproof.yaml cannot name a custom evaluator yet.
- A judge sees text: JSON of the input, reference and output. It does not see images or audio.
- Twenty labels give a wide agreement interval. Label more, at random and blind, for a judge you rely on.
- A local sample is good-faith only. For a judge other people rely on, push the run and let a hosted workspace draw the sample (Judges).