Skip to content

Guides

Tutorial: an application behind an HTTP endpoint

Evaluate a service you reach over HTTP, without importing its code: point Oloproof at the URL, run the suite, find what it misses, deploy a change, and compare. A tiny local service stands in for yours, so everything runs offline.

What you will build

An order-intake service that reads a customer message and extracts two fields: an intent (where_is_order, cancel, return or other) and an order_id (four digits, or null). You will check the format of every response, which needs no reference, and measure whether each field is right, which does. Terms such as case, run, metric and gate are defined in Concepts.

Prerequisites

  • Python 3.11 or later, and Oloproof in a virtual environment:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • The example project, which ships with the package. Copy it into a new directory and work there (httpx, which client.py imports, is installed with Oloproof):
oloproof init --example http order-intake
cd order-intake
  • Port 8765 free on this machine. If it is taken, pick another and change it in both the server command and oloproof.yaml.

Nothing here uses a provider, an API key or the internet: the service listens on 127.0.0.1.

The files

order-intake/
  server.py             the stand-in service (Python standard library only)
  client.py             a callable that calls the service with a token (used near the end)
  oloproof.yaml         the suite: dataset, HTTP system, evaluators
  release.yaml          rules for a run
  compare.yaml          a rule for a comparison
  data/messages.jsonl   20 cases

Run the commands from order-intake/. Start the service in a second terminal and leave it running:

python server.py --port 8765
intake service (v1) on http://127.0.0.1:8765/extract

The HTTP contract

For each case Oloproof sends one request: the case's input as the JSON body, with the method you declare (POST by default). It reads the response as JSON. A status of 400 or above, a timeout or a refused connection is recorded as an execution error for that case, never as a wrong answer.

Request and response for one case:

POST /extract
{"message": "Where is order 1042? It has not arrived."}

200 OK
{"result": {"intent": "where_is_order", "order_id": "1042"}, "service": {"rules": "v1"}}

output_path: result tells Oloproof to keep only result as the case's output; without it the whole body is the output. A dotted path such as data.answer reaches deeper.

version: 1
project: order-intake
dataset: data/messages.jsonl
system:
  name: order-intake
  http:
    url: http://127.0.0.1:8765/extract
    method: POST
    version: rules-v1
    output_path: result
    timeout_s: 30
evaluators:
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [intent, order_id]
      properties:
        intent: {enum: [where_is_order, cancel, return, other]}
        order_id: {type: [string, "null"], pattern: '^\d{4}$'}
      additionalProperties: false
  - type: exact_match
    criterion: intent_correct
    field: intent
  - type: exact_match
    criterion: order_id_correct
    field: order_id

An HTTP system must declare a version. Oloproof cannot see a deployment: it caches each case's output under the URL, the method, the output path and that version, so the version is how you tell it the service changed. Forget to change it and a new deployment is never called.

Headers and authentication cannot be configured on system.http yet. The section on tokens below shows the workaround.

The dataset

{"id":"m04","input":{"message":"Has order #5120 shipped yet?"},"expected":{"intent":"where_is_order","order_id":"5120"}}
{"id":"m07","input":{"message":"Do you ship to Canada?"},"expected":{"intent":"other","order_id":null}}
{"id":"m16","input":{"message":"Please refund and take back the lamp from order #1560."},"expected":{"intent":"return","order_id":"1560"}}

input is the request body exactly. expected holds the reference for each field; null is a real reference value, meaning "there is no order id in this message".

Choosing the evaluators

CriterionEvaluatorNeeds expectedMeasures
format_validjson_schema over the outputnoformat: a known intent and a well-formed id
intent_correctexact_match on intentyestask success for the first field
order_id_correctexact_match on order_idyestask success for the second field

The schema would pass {"intent": "other", "order_id": null} for every message: well formed and useless. Only the reference checks say whether the service did its job. Scoring the fields separately shows which one fails, which one combined check would hide.

release.yaml allows no format failure and asks for each field to be right at least 70% of the time, judged on the 95% interval:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: intent-floor
    metric: intent_correct
    min: 0.70
  - id: order-id-floor
    metric: order_id_correct
    min: 0.70

Run it

oloproof run
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format   │ format_valid     │ PASS                  │ observed_failures_within_limit │
│ intent-floor   │ intent_correct   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ order-id-floor │ order_id_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ format_valid     │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ intent_correct   │ 75.0%    │ [50.8%, 91.4%]  │ 15 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 75.0%    │ [50.8%, 91.4%]  │ 15 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 miss

Every response is well formed. Each field is right 15 times in 20, but 20 cases leave an interval down to 50.8%, so neither floor of 70% is shown: INSUFFICIENT_EVIDENCE, and the gate blocks with exit 3.

Inspect the failures

oloproof inspect RUN_ID --failures
8 of 20 cases failed, errored or did not finish

m04
  output: {"intent": "where_is_order", "order_id": null}
  order_id_correct: failed

m06
  output: {"intent": "other", "order_id": "7011"}
  intent_correct: failed
...
m18
  output: {"intent": "other", "order_id": null}
  intent_correct: failed
  order_id_correct: failed

Read the inputs beside them (oloproof inspect RUN_ID --case m04) and two faults appear: the id pattern only matches "order 1234", not "order #5120", "order no. 8123" or a bare "#1673"; and phrasings such as "send back", "stop order" and "where's my parcel" map to no intent. That is the next action: widen both rules.

Deploy a change

Stop the service and start the candidate, which carries those fixes:

python server.py --port 8765 --rules v2

Change version: rules-v1 to version: rules-v2 in oloproof.yaml, because the URL did not change and Oloproof would otherwise reuse the old outputs. Then:

oloproof run
Gate: ALLOW (exit 0)
│ intent_correct   │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │

Both floors pass and the command exits 0.

Compare the two deployments

compare.yaml asks whether the candidate's intents are better than the baseline's:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: intent-better
    kind: superiority
    metric: intent_correct
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
format_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
intent_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
order_id_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
Decisions
  intent-better  intent_correct  superiority  INSUFFICIENT_EVIDENCE  interval_overlaps_zero
Gate: BLOCK (exit 3)

The candidate passes its own floors, yet the comparison cannot show it is better: five cases changed, and over 20 paired cases the interval for the gain still includes zero. Both statements are true at once. "Meets the requirement" and "beats the baseline" are separate questions, and a suite this small answers the second only for large effects. More real messages is the remedy.

A service that needs a token

Start the service so that it demands a bearer token:

INTAKE_TOKEN=s3cret python server.py --port 8765 --rules v2 --require-token

system.http sends no custom headers, so with a new version (say rules-v2-auth) every call is refused:

Gate: BLOCK (exit 3)
│ valid-format   │ format_valid     │ INSUFFICIENT_EVIDENCE │ no_observations │
│ intent_correct   │          │ [0.0%, 100.0%] │ 0 / 0 observed · 20 missing · 0 excluded │

and oloproof inspect RUN_ID --failures shows execution ERROR: TransientError: system returned HTTP 401 on each case. Execution errors are missing evidence, not failures: nothing was observed, so every rule is INSUFFICIENT_EVIDENCE.

The workaround is a Python callable that makes the request itself. client.py adds the header from an environment variable, so the token never enters oloproof.yaml or any stored record:

URL = os.environ.get("INTAKE_URL", "http://127.0.0.1:8765/extract")


def extract(case: dict[str, Any]) -> dict[str, Any]:
    headers = {"Authorization": f"Bearer {os.environ['INTAKE_TOKEN']}"}
    response = httpx.post(URL, json=case, headers=headers, timeout=30)
    response.raise_for_status()
    return response.json()["result"]

Replace the http: block in oloproof.yaml with:

system:
  name: order-intake
  version: rules-v2
  callable: client:extract

and run with the token in the environment:

INTAKE_TOKEN=s3cret oloproof run

All 20 cases are observed again and the gate allows. The callable's cache follows its own source and version, not the service behind it, so the same rule applies: change version when you deploy.

Troubleshooting

SymptomCause and fix
system connection failed (ConnectError) on every caseThe service is not running, or listens on another port.
system connection failed (RemoteProtocolError)Something else answers on that port. Choose a free one.
KeyError: "missing output path 'results'"output_path names a field the response does not have.
system returned HTTP 401 or 403The endpoint needs credentials: use the callable workaround.
You deployed a change and the cache line reads all hitsThe version did not change, so the stored outputs were reused.
an HTTP system needs a declared versionAdd version under system or system.http.

Limitations

  • No custom headers, authentication, query parameters or request templating on system.http: the case input is the JSON body as it stands. Use a callable for anything else.
  • Responses must be JSON. Streaming responses are not read as a stream.
  • Oloproof cannot detect a deployment; the declared version is the whole identity of what answered.
  • Your service owns its state: Oloproof sends requests and records answers; it does not reset, sandbox or roll back anything a request changes.