Skip to content

Guides

Tutorial: text generation with a rubric judge

Evaluate a function that writes free text, here a ticket summariser, with format checks and a rubric judge; measure that judge against a person's labels before it may decide anything; then compare a real change. The judge runs on this machine with no model and no network, and an optional step swaps in a real model.

What you will build

A summariser that turns a support ticket into one or two sentences. "Good" is a judgment, not a string match, so task success is decided by an LLM judge with a rubric: does the summary state the facts an agent needs? Two deterministic evaluators check format, which needs no reference. Terms such as case, run, metric, judge and gate are defined in Concepts.

The same shape fits extraction or any other generation: a function returns text in a dictionary, the reference says what a good answer must contain, and a rubric says how to decide.

Prerequisites

  • Python 3.11 or later, and Oloproof in a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • The example project, which ships with the package. Copy it into a new directory and work there:
oloproof init --example generation ticket-summaries
cd ticket-summaries
  • Port 8799 free for the stand-in judge (change it in both places if not).

Every step until "Optional: a real model as the judge" is offline and deterministic: no API key, no provider account, no cost.

The files

ticket-summaries/
  app.py                        the summariser under test (baseline)
  app_v2.py                     the candidate change
  judge_server.py               a stand-in judge speaking the OpenAI API on 127.0.0.1
  rubrics/covers_facts.md       the judge's rubric
  oloproof.yaml                 the suite
  release.yaml                  rules for a run
  compare.yaml                  a rule for a comparison
  data/tickets.jsonl            20 cases
  labels/reviewer_verdicts.csv  one person's verdicts on the baseline's summaries
  fill_labels.py                copies those verdicts into a labelling sheet

Run every command from ticket-summaries/.

The stand-in judge, and what it is not

A rubric judge is an evaluator that sends a prompt (the rubric, the case's input, its expected and the output) to a model and reads back {"pass": true|false, "rationale": "..."}. Oloproof talks to any server speaking the OpenAI chat API, and a server on localhost needs no key.

judge_server.py is such a server, but it is not a model. It passes a summary only when it contains every phrase under must_mention in the case's expected, ignoring case. That is a fixed rule, so the tutorial gives the same numbers on every machine. It cannot notice an invented fact, which a real model judge is asked to. Start it in a second terminal and leave it running:

python judge_server.py --port 8799
stand-in judge on http://127.0.0.1:8799/v1

The application and its adapter

# app.py
@system(name="ticket-summariser", version="first-sentence")
def summarise(case: dict[str, Any]) -> dict[str, str]:
    return {"summary": sentences(str(case["ticket"]))[0]}

The adapter for a Python application is the function: it receives the case's input and returns a dictionary. For your own generator, call your model or chain inside it and return the text under a key. Oloproof calls it once per case and caches the output on the function's source and declared version; it does not manage your model client, prompts or state. List files the function reads, such as a prompt template, under system.code_paths.

The dataset

{"id":"t01","input":{"ticket":"Hello. Order 1042 arrived with a cracked screen. I would like a replacement, not a refund."},"expected":{"must_mention":["1042","cracked","replacement"]}}
{"id":"t06","input":{"ticket":"Please cancel my subscription at the end of this month. I am moving abroad."},"expected":{"must_mention":["cancel","end of this month"]}}

input is what the function receives. expected is the reference the judge reads: here a list of facts the summary must carry, not a full reference summary, because many different summaries are correct. The output for t01 is {"summary": "Hello."}.

Choosing the evaluators

version: 1
project: ticket-summaries
dataset: data/tickets.jsonl
system:
  name: ticket-summariser
  version: first-sentence
  callable: app:summarise
  timeout_s: 30
evaluators:
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [summary]
      properties:
        summary: {type: string, minLength: 1}
      additionalProperties: false
  - type: regex
    criterion: short_enough
    field: summary
    pattern: '^.{1,160}$'
    pass_if: match
  - type: rubric_judge
    criterion: covers_facts
    provider: openai_compatible
    model: stand-in-judge
    base_url: http://127.0.0.1:8799/v1
    rubric_file: rubrics/covers_facts.md
CriterionEvaluatorNeeds expectedMeasures
format_validjson_schemanoformat: one non-empty string field
short_enoughregexnoformat: at most 160 characters
covers_factsrubric_judgeyestask success, as the rubric defines it

Hello. passes both format checks. Only the judge says it is a useless summary. A judge can also run without a reference: a rubric such as "PASS if the summary contains no greeting" reads only the input and output, and a case with no expected is still judged. What it cannot do then is check facts against an answer you trust.

The rubric:

PASS when the summary states every fact listed under must_mention in the expected answer, in
words a support agent would recognise, and adds nothing the ticket does not say.
FAIL when any listed fact is missing, changed or contradicted.

The policy

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: short-enough
    metric: short_enough
    kind: observed_count
    max_failures: 0
  - id: covers-facts-floor
    metric: covers_facts
    min: 0.60

require_validated_evaluators: true is the engine's default, written out here because it is the point of this tutorial: a judge nobody has compared with people may not decide a rule.

Run it

oloproof run
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format       │ format_valid │ PASS                  │ observed_failures_within_limit │
│ short-enough       │ short_enough │ PASS                  │ observed_failures_within_limit │
│ covers-facts-floor │ covers_facts │ INSUFFICIENT_EVIDENCE │ evaluator_not_validated        │
covers-facts-floor: the judge (or model or custom evaluator) behind this rule has not been measured against
people yet, so it may not decide.
  Label a sample:  oloproof review run_01M4... --criterion covers_facts --by YOU --sample 20
  Then measure it: oloproof evaluators validate EVALUATOR_ID --by YOU (ids: oloproof evaluators list)
│ format_valid │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ short_enough │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ covers_facts │ 45.0%    │ [23.0%, 68.5%]  │ 9 / 20 observed · 0 missing · 0 excluded  │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 miss

The format rules pass. The judge passed 9 of 20 summaries, but the rule is INSUFFICIENT_EVIDENCE with the reason evaluator_not_validated, and the gate blocks with exit 3. The rule did not decide on the 45%: a judge's error rate is unknown until it is measured, so an interval built on its verdicts would carry an unstated error. The engine reports this as INSUFFICIENT_EVIDENCE, not MANUAL_REVIEW or FAIL: the evidence to decide is missing, and the output prints the two commands that supply it.

Inspect the failures

oloproof inspect RUN_ID --failures
11 of 20 cases failed, errored or did not finish

t01
  output: {"summary": "Hello."}
  covers_facts: failed
    judge text, not verified: missing: 1042, cracked, replacement

t02
  output: {"summary": "I was charged twice for order 2210."}
  covers_facts: failed
    judge text, not verified: missing: 49
...

The judge's rationale is shown as "judge text, not verified": it is the model's explanation, not evidence. The pattern is clear anyway: the first sentence is often a greeting.

Measure the judge against a person

Validation compares the judge's verdicts with a person's on the same answers. Draw a random sample of the run's cases into a sheet. The judge's verdicts are left out of it, so the labeller is not anchored on them:

oloproof labels export RUN_ID --criterion covers_facts --sample 20 --local --out sample.csv
Wrote 20 cases to sample.csv, drawn at random with seed 2701013296, without the judge's verdict.
  This is a local sample, good-faith only, because it was drawn on this machine.
Fill in `passed` (pass or fail) and `labelled_by` on each row you judge, then run `oloproof labels import sample.csv`.

--local draws on this machine without asking a hosted workspace; the engine chooses the seed. With 20 cases, a sample of 20 is all of them. In practice, a person reads each row's ticket and summary and fills in passed. For this tutorial, labels/reviewer_verdicts.csv holds verdicts a reviewer gave on the baseline's summaries, and fill_labels.py copies them into the sheet:

python fill_labels.py sample.csv
oloproof labels import sample.csv
filled 20 rows of sample.csv
Recorded 20 labels from sample.csv (20 measurement).

The reviewer disagreed with the judge once: on t02 ("I was charged twice for order 2210.") they judged the missing amount immaterial and passed it. Labels name the exact answer they judged, so these verdicts apply to the baseline run only.

Find the judge's version id and validate it:

oloproof evaluators list
oloproof evaluators validate EVALUATOR_ID --by alice
covers_facts  LLM_JUDGE  UNVALIDATED  (declared)  sha256:a662...

covers_facts: sha256:a662... is now VALIDATED
  agreement 95.0% [75.1%, 99.9%] · 19 of 20 labelled cases agreed · 0 labelled but not judged · kappa 0.900
  bias -5.0 points [-32.4, +20.7] · the judge's pass rate minus the people's · 20 cases · 0 labelled but not judged
  passes what people pass 90.0% [55.4%, 99.8%] · the judge passed 9 of 10 cases people passed · 0 labelled but not judged
  fails what people fail 100.0% [69.1%, 100.0%] · the judge failed 10 of 10 cases people failed · 0 labelled but not judged

Read the intervals, not the 95%: 20 labels show agreement of at least 75.1%. A policy can demand more with minimum_evaluator_agreement, which compares that lower bound, and validate refuses a judge below it. The Judges guide covers the bar, bias, probes and oloproof review for labelling in the terminal.

Now re-decide the stored run without calling the summariser or the judge:

oloproof gate RUN_ID --policy release.yaml
valid-format: PASS (observed_failures_within_limit)
short-enough: PASS (observed_failures_within_limit)
covers-facts-floor: INSUFFICIENT_EVIDENCE (interval_overlaps_threshold)
  no sample size would make this PASS: the observed rate (0.500) is itself below the threshold (0.600), so more cases would move it toward FAIL
Gate: BLOCK (exit 3)

The judge may decide now, and the decision is about the summariser: the rate it quotes, 0.500, is not the judge's 45%. Because this run has a blind, random sample of measurement labels, the gate reads the judge corrected by those labels ("Judge-corrected gates" in the Judges guide). The correction is PPI, prediction-powered inference: it uses the labelled sample to measure how far the judge's rate sits from the people's, and moves the estimate and widens the interval by that much. That is also what the export's notes about PPI refer to. Either way, the baseline does not meet the floor, and more cases would not change that.

Make a real change

app_v2.py skips short pleasantries and keeps the next two sentences. Copy it over app.py, set version: skip-pleasantries under system in oloproof.yaml, keep the judge running, and:

oloproof run
Gate: ALLOW (exit 0)
│ covers-facts-floor │ covers_facts │ PASS  │ lower_bound_meets_minimum      │
│ covers_facts │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 6 hit/54 miss

The judge is the same validated version, so its rule decides directly. Six judgments came from the cache, on summaries both versions wrote identically. Nobody labelled these new summaries; the judge's validation is what lets its verdicts stand.

Compare the candidate with the baseline

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
require_validated_evaluators: true
rules:
  - id: covers-more-facts
    kind: superiority
    metric: covers_facts
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
format_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
short_enough: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
covers_facts: +55.0 points [+13.0, +84.4] · 20 paired · 0 missing · 0 excluded
Decisions
  covers-more-facts  covers_facts  superiority  PASS  difference_above_zero
Gate: ALLOW (exit 0)

A comparison does not apply the PPI correction: it compares the judge's own verdicts on the two runs, which is why the gain starts from the judge's 45% and not from the corrected 0.500 above. Eleven summaries improved and none got worse; the interval for the gain is entirely above zero, so the superiority rule passes and the command exits 0. Format is guarded by the run rules, which allow no failure, rather than by a comparison: over 20 cases a comparison of two perfect format scores could only say the difference is within 23.6 points.

Optional: a real model as the judge

This step leaves the offline path. It needs a model server, and with a cloud provider a key and money.

  • Local, no key and no cost: Ollama, LM Studio or llama.cpp on localhost. Pull a chat model (for Ollama, ollama pull llama3.1).
  • Cloud: provider: anthropic or openai with api_key_env naming the variable that holds your key, or openai_compatible with base_url and api_key_env. Each case is one judge call (two when the first reply is not valid JSON), billed at your provider's rates, and Oloproof never calls a judge again for an answer it has already judged.

Write the draft judge in a file of its own, as it would appear under evaluators::

# live_judge.yaml
type: rubric_judge
criterion: covers_facts
provider: openai_compatible
model: llama3.1
base_url: http://localhost:11434/v1
rubric_file: rubrics/covers_facts.md

and try it against the answers your reviewer already labelled, without validating or adopting it:

oloproof evaluators try live_judge.yaml

Local servers answer one request at a time by default; add concurrency: {system: 2, judge: 2} to oloproof.yaml so queued calls do not time out. A run of this step with a small local model (qwen2.5vl) on a laptop printed:

covers_facts: draft sha256:b88a... on 20 labelled cases · 20 judged now, 0 from cache, 11 errored
  agreement 88.9% [19.1%, 99.9%] · 8 of 9 labelled cases agreed · 11 labelled but not judged · kappa 0.769

Eleven calls timed out, and the agreement interval counts each one both ways, so it reaches down to 19.1%: a judge that does not answer is not measured. A larger model, a longer timeout, or fewer concurrent calls is the fix. To adopt the model, put it in oloproof.yaml in place of the stand-in. That is a new evaluator version: its configuration (model, endpoint, rubric) is its identity, so the stand-in's validation does not carry over. Run the baseline again with it and validate it against the labels, as above.

Troubleshooting

SymptomCause and fix
covers_facts all missing, no_observationsThe judge server is not running or not on base_url. Every judge call errored; oloproof inspect RUN_ID --failures shows why.
evaluator_not_validated after you validatedYou changed the judge (model, endpoint, port, rubric) and made a new version. Validate that one.
labels import refuses the file and names a rowThe row names a case or execution the run does not hold; export again from the run you label.
labels export says a workspace could not be reachedYou are logged in to one, so it asked it to draw. --local draws here instead.
A cloud judge fails before any callIts key is not in the variable api_key_env names.

Limitations

  • The stand-in judge is a phrase match. It demonstrates the workflow, not judging quality.
  • There are no BLEU, ROUGE or embedding-similarity evaluators. In the SDK, write one with @evaluator; oloproof.yaml cannot name a custom evaluator yet.
  • A judge sees text: JSON of the input, reference and output. It does not see images or audio.
  • Twenty labels give a wide agreement interval. Label more, at random and blind, for a judge you rely on.
  • A local sample is good-faith only. For a judge other people rely on, push the run and let a hosted workspace draw the sample (Judges).