Guides
Tutorial: an application behind an HTTP endpoint
Evaluate a service you reach over HTTP, without importing its code: point Oloproof at the URL, run the suite, find what it misses, deploy a change, and compare. A tiny local service stands in for yours, so everything runs offline.
What you will build
An order-intake service that reads a customer message and extracts two fields: an intent (where_is_order, cancel, return or other) and an order_id (four digits, or null). You will check the format of every response, which needs no reference, and measure whether each field is right, which does. Terms such as case, run, metric and gate are defined in Concepts.
Prerequisites
- Python 3.11 or later, and Oloproof in a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof- The example project, which ships with the package. Copy it into a new directory and work there (httpx, which client.py imports, is installed with Oloproof):
oloproof init --example http order-intake
cd order-intake- Port 8765 free on this machine. If it is taken, pick another and change it in both the server command and oloproof.yaml.
Nothing here uses a provider, an API key or the internet: the service listens on 127.0.0.1.
The files
order-intake/
server.py the stand-in service (Python standard library only)
client.py a callable that calls the service with a token (used near the end)
oloproof.yaml the suite: dataset, HTTP system, evaluators
release.yaml rules for a run
compare.yaml a rule for a comparison
data/messages.jsonl 20 casesRun the commands from order-intake/. Start the service in a second terminal and leave it running:
python server.py --port 8765intake service (v1) on http://127.0.0.1:8765/extractThe HTTP contract
For each case Oloproof sends one request: the case's input as the JSON body, with the method you declare (POST by default). It reads the response as JSON. A status of 400 or above, a timeout or a refused connection is recorded as an execution error for that case, never as a wrong answer.
Request and response for one case:
POST /extract
{"message": "Where is order 1042? It has not arrived."}
200 OK
{"result": {"intent": "where_is_order", "order_id": "1042"}, "service": {"rules": "v1"}}output_path: result tells Oloproof to keep only result as the case's output; without it the whole body is the output. A dotted path such as data.answer reaches deeper.
version: 1
project: order-intake
dataset: data/messages.jsonl
system:
name: order-intake
http:
url: http://127.0.0.1:8765/extract
method: POST
version: rules-v1
output_path: result
timeout_s: 30
evaluators:
- type: json_schema
criterion: format_valid
field: null
schema:
type: object
required: [intent, order_id]
properties:
intent: {enum: [where_is_order, cancel, return, other]}
order_id: {type: [string, "null"], pattern: '^\d{4}$'}
additionalProperties: false
- type: exact_match
criterion: intent_correct
field: intent
- type: exact_match
criterion: order_id_correct
field: order_idAn HTTP system must declare a version. Oloproof cannot see a deployment: it caches each case's output under the URL, the method, the output path and that version, so the version is how you tell it the service changed. Forget to change it and a new deployment is never called.
Headers and authentication cannot be configured on system.http yet. The section on tokens below shows the workaround.
The dataset
{"id":"m04","input":{"message":"Has order #5120 shipped yet?"},"expected":{"intent":"where_is_order","order_id":"5120"}}
{"id":"m07","input":{"message":"Do you ship to Canada?"},"expected":{"intent":"other","order_id":null}}
{"id":"m16","input":{"message":"Please refund and take back the lamp from order #1560."},"expected":{"intent":"return","order_id":"1560"}}input is the request body exactly. expected holds the reference for each field; null is a real reference value, meaning "there is no order id in this message".
Choosing the evaluators
| Criterion | Evaluator | Needs expected | Measures |
|---|---|---|---|
| format_valid | json_schema over the output | no | format: a known intent and a well-formed id |
| intent_correct | exact_match on intent | yes | task success for the first field |
| order_id_correct | exact_match on order_id | yes | task success for the second field |
The schema would pass {"intent": "other", "order_id": null} for every message: well formed and useless. Only the reference checks say whether the service did its job. Scoring the fields separately shows which one fails, which one combined check would hide.
release.yaml allows no format failure and asks for each field to be right at least 70% of the time, judged on the 95% interval:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: valid-format
metric: format_valid
kind: observed_count
max_failures: 0
- id: intent-floor
metric: intent_correct
min: 0.70
- id: order-id-floor
metric: order_id_correct
min: 0.70Run it
oloproof runRun run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format │ format_valid │ PASS │ observed_failures_within_limit │
│ intent-floor │ intent_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ order-id-floor │ order_id_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold │
│ format_valid │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ intent_correct │ 75.0% │ [50.8%, 91.4%] │ 15 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 75.0% │ [50.8%, 91.4%] │ 15 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 missEvery response is well formed. Each field is right 15 times in 20, but 20 cases leave an interval down to 50.8%, so neither floor of 70% is shown: INSUFFICIENT_EVIDENCE, and the gate blocks with exit 3.
Inspect the failures
oloproof inspect RUN_ID --failures8 of 20 cases failed, errored or did not finish
m04
output: {"intent": "where_is_order", "order_id": null}
order_id_correct: failed
m06
output: {"intent": "other", "order_id": "7011"}
intent_correct: failed
...
m18
output: {"intent": "other", "order_id": null}
intent_correct: failed
order_id_correct: failedRead the inputs beside them (oloproof inspect RUN_ID --case m04) and two faults appear: the id pattern only matches "order 1234", not "order #5120", "order no. 8123" or a bare "#1673"; and phrasings such as "send back", "stop order" and "where's my parcel" map to no intent. That is the next action: widen both rules.
Deploy a change
Stop the service and start the candidate, which carries those fixes:
python server.py --port 8765 --rules v2Change version: rules-v1 to version: rules-v2 in oloproof.yaml, because the URL did not change and Oloproof would otherwise reuse the old outputs. Then:
oloproof runGate: ALLOW (exit 0)
│ intent_correct │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 100.0% │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │Both floors pass and the command exits 0.
Compare the two deployments
compare.yaml asks whether the candidate's intents are better than the baseline's:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- id: intent-better
kind: superiority
metric: intent_correctoloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yamlformat_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
intent_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
order_id_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
Decisions
intent-better intent_correct superiority INSUFFICIENT_EVIDENCE interval_overlaps_zero
Gate: BLOCK (exit 3)The candidate passes its own floors, yet the comparison cannot show it is better: five cases changed, and over 20 paired cases the interval for the gain still includes zero. Both statements are true at once. "Meets the requirement" and "beats the baseline" are separate questions, and a suite this small answers the second only for large effects. More real messages is the remedy.
A service that needs a token
Start the service so that it demands a bearer token:
INTAKE_TOKEN=s3cret python server.py --port 8765 --rules v2 --require-tokensystem.http sends no custom headers, so with a new version (say rules-v2-auth) every call is refused:
Gate: BLOCK (exit 3)
│ valid-format │ format_valid │ INSUFFICIENT_EVIDENCE │ no_observations │
│ intent_correct │ │ [0.0%, 100.0%] │ 0 / 0 observed · 20 missing · 0 excluded │and oloproof inspect RUN_ID --failures shows execution ERROR: TransientError: system returned HTTP 401 on each case. Execution errors are missing evidence, not failures: nothing was observed, so every rule is INSUFFICIENT_EVIDENCE.
The workaround is a Python callable that makes the request itself. client.py adds the header from an environment variable, so the token never enters oloproof.yaml or any stored record:
URL = os.environ.get("INTAKE_URL", "http://127.0.0.1:8765/extract")
def extract(case: dict[str, Any]) -> dict[str, Any]:
headers = {"Authorization": f"Bearer {os.environ['INTAKE_TOKEN']}"}
response = httpx.post(URL, json=case, headers=headers, timeout=30)
response.raise_for_status()
return response.json()["result"]Replace the http: block in oloproof.yaml with:
system:
name: order-intake
version: rules-v2
callable: client:extractand run with the token in the environment:
INTAKE_TOKEN=s3cret oloproof runAll 20 cases are observed again and the gate allows. The callable's cache follows its own source and version, not the service behind it, so the same rule applies: change version when you deploy.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| system connection failed (ConnectError) on every case | The service is not running, or listens on another port. |
| system connection failed (RemoteProtocolError) | Something else answers on that port. Choose a free one. |
| KeyError: "missing output path 'results'" | output_path names a field the response does not have. |
| system returned HTTP 401 or 403 | The endpoint needs credentials: use the callable workaround. |
| You deployed a change and the cache line reads all hits | The version did not change, so the stored outputs were reused. |
| an HTTP system needs a declared version | Add version under system or system.http. |
Limitations
- No custom headers, authentication, query parameters or request templating on system.http: the case input is the JSON body as it stands. Use a callable for anything else.
- Responses must be JSON. Streaming responses are not read as a stream.
- Oloproof cannot detect a deployment; the declared version is the whole identity of what answered.
- Your service owns its state: Oloproof sends requests and records answers; it does not reset, sandbox or roll back anything a request changes.