跳到主要内容

指南

教程:HTTP 端点背后的应用

评估一个通过 HTTP 访问的服务,而不导入它的代码:把 Oloproof 指向 URL,运行测试套件,找出它的遗漏,部署一次变更,然后进行比较。一个小型本地服务代替你的服务,所以一切都可以离线运行。

你将构建什么

一个订单受理服务,它读取一条客户消息并抽取两个字段:一个 intent(where_is_order、cancel、return 或 other)和一个 order_id(四位数字,或 null)。你将检查每个响应的格式,这不需要参考;并度量每个字段是否正确,这需要参考。用例、运行、指标和门禁等术语在核心概念中有定义。

前提条件

  • Python 3.11 或更高版本,以及安装在虚拟环境中的 Oloproof:
python3 -m venv .venv
. .venv/bin/activate
pip install oloproof
  • 示例项目,它随软件包一起提供。把它复制到一个新目录并在那里工作(client.py 导入的 httpx 会随 Oloproof 一起安装):
oloproof init --example http order-intake
cd order-intake
  • 本机的 8765 端口空闲。如果已被占用,另选一个,并在服务器命令和 oloproof.yaml 中同时修改。

这里没有任何内容使用提供方、API 密钥或互联网:服务监听在 127.0.0.1 上。

文件

order-intake/
  server.py             the stand-in service (Python standard library only)
  client.py             a callable that calls the service with a token (used near the end)
  oloproof.yaml         the suite: dataset, HTTP system, evaluators
  release.yaml          rules for a run
  compare.yaml          a rule for a comparison
  data/messages.jsonl   20 cases

在 order-intake/ 中运行这些命令。在第二个终端中启动服务并让它保持运行:

python server.py --port 8765
intake service (v1) on http://127.0.0.1:8765/extract

HTTP 契约

对于每个用例,Oloproof 发送一个请求:以该用例的 input 作为 JSON 请求体,使用你声明的方法(默认为 POST)。它把响应作为 JSON 读取。400 或以上的状态码、超时或连接被拒绝,都会被记录为该用例的执行错误,绝不会被记为错误答案。

一个用例的请求和响应:

POST /extract
{"message": "Where is order 1042? It has not arrived."}

200 OK
{"result": {"intent": "where_is_order", "order_id": "1042"}, "service": {"rules": "v1"}}

output_path: result 告诉 Oloproof 只把 result 保留为该用例的输出;没有它时,整个响应体就是输出。像 data.answer 这样的点分路径可以深入到更里层。

version: 1
project: order-intake
dataset: data/messages.jsonl
system:
  name: order-intake
  http:
    url: http://127.0.0.1:8765/extract
    method: POST
    version: rules-v1
    output_path: result
    timeout_s: 30
evaluators:
  - type: json_schema
    criterion: format_valid
    field: null
    schema:
      type: object
      required: [intent, order_id]
      properties:
        intent: {enum: [where_is_order, cancel, return, other]}
        order_id: {type: [string, "null"], pattern: '^\d{4}$'}
      additionalProperties: false
  - type: exact_match
    criterion: intent_correct
    field: intent
  - type: exact_match
    criterion: order_id_correct
    field: order_id

HTTP 系统必须声明一个 version。Oloproof 看不到部署:它按 URL、方法、输出路径和这个版本缓存每个用例的输出,所以版本就是你告诉它服务已经变化的方式。忘了修改它,新的部署就永远不会被调用。

目前还无法在 system.http 上配置请求头和身份验证。下面关于令牌的一节展示了变通办法。

数据集

{"id":"m04","input":{"message":"Has order #5120 shipped yet?"},"expected":{"intent":"where_is_order","order_id":"5120"}}
{"id":"m07","input":{"message":"Do you ship to Canada?"},"expected":{"intent":"other","order_id":null}}
{"id":"m16","input":{"message":"Please refund and take back the lamp from order #1560."},"expected":{"intent":"return","order_id":"1560"}}

input 就是请求体本身。expected 保存每个字段的参考;null 是一个真正的参考值,意思是“这条消息中没有订单 id”。

选择评估器

判据评估器需要 expected度量
format_valid对输出使用 json_schema否格式:一个已知的意图和一个格式正确的 id
intent_correct对 intent 使用 exact_match是第一个字段的任务成功
order_id_correct对 order_id 使用 exact_match是第二个字段的任务成功

对每条消息都返回 {"intent": "other", "order_id": null},模式检查也会通过:格式正确,却毫无用处。只有参考检查才能说明服务是否完成了它的工作。分别为各字段打分可以显示是哪一个失败了,而一个合并的检查会把这一点掩盖掉。

release.yaml 不允许任何格式失败,并要求每个字段至少在 70% 的情况下正确,依据 95% 区间来判断:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: valid-format
    metric: format_valid
    kind: observed_count
    max_failures: 0
  - id: intent-floor
    metric: intent_correct
    min: 0.70
  - id: order-id-floor
    metric: order_id_correct
    min: 0.70

运行

oloproof run
Run run_01M4... [DECIDED/COMPLETE]
Gate: BLOCK (exit 3)
│ valid-format   │ format_valid     │ PASS                  │ observed_failures_within_limit │
│ intent-floor   │ intent_correct   │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ order-id-floor │ order_id_correct │ INSUFFICIENT_EVIDENCE │ interval_overlaps_threshold    │
│ format_valid     │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ intent_correct   │ 75.0%    │ [50.8%, 91.4%]  │ 15 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 75.0%    │ [50.8%, 91.4%]  │ 15 / 20 observed · 0 missing · 0 excluded │
Cache: execution 0 hit/20 miss; judgment 0 hit/60 miss

每个响应的格式都正确。每个字段在 20 次中正确 15 次,但 20 个用例留下的区间向下延伸到 50.8%,所以两个 70% 的下限都无法证明:INSUFFICIENT_EVIDENCE,门禁以退出码 3 阻止。

查看失败项

oloproof inspect RUN_ID --failures
8 of 20 cases failed, errored or did not finish

m04
  output: {"intent": "where_is_order", "order_id": null}
  order_id_correct: failed

m06
  output: {"intent": "other", "order_id": "7011"}
  intent_correct: failed
...
m18
  output: {"intent": "other", "order_id": null}
  intent_correct: failed
  order_id_correct: failed

把它们旁边的输入也读一读(oloproof inspect RUN_ID --case m04),就会出现两个缺陷:id 模式只匹配“order 1234”,而不匹配“order #5120”、“order no. 8123”或单独的“#1673”;而像“send back”、“stop order”和“where's my parcel”这样的说法映射不到任何意图。这就是下一步:放宽这两条规则。

部署一次变更

停止服务并启动候选版本,它带有这些修复:

python server.py --port 8765 --rules v2

在 oloproof.yaml 中把 version: rules-v1 改为 version: rules-v2,因为 URL 没有变化,否则 Oloproof 会复用旧的输出。然后:

oloproof run
Gate: ALLOW (exit 0)
│ intent_correct   │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │
│ order_id_correct │ 100.0%   │ [83.1%, 100.0%] │ 20 / 20 observed · 0 missing · 0 excluded │

两个下限都通过,命令以 0 退出。

比较两次部署

compare.yaml 问的是候选的意图是否比基线的更好:

version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
  - id: intent-better
    kind: superiority
    metric: intent_correct
oloproof compare CANDIDATE_RUN_ID BASELINE_RUN_ID --policy compare.yaml
format_valid: +0.0 points [-23.6, +23.6] · 20 paired · 0 missing · 0 excluded
intent_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
order_id_correct: +25.0 points [-8.5, +58.3] · 20 paired · 0 missing · 0 excluded
Decisions
  intent-better  intent_correct  superiority  INSUFFICIENT_EVIDENCE  interval_overlaps_zero
Gate: BLOCK (exit 3)

候选通过了它自己的下限,但比较无法证明它更好:有五个用例发生了变化,而在 20 个配对用例上,提升的区间仍然包含零。两种说法同时成立。“满足要求”和“胜过基线”是两个不同的问题,而这么小的测试套件只能对很大的效应回答第二个问题。更多真实消息才是解决办法。

需要令牌的服务

启动服务,让它要求一个 bearer 令牌:

INTAKE_TOKEN=s3cret python server.py --port 8765 --rules v2 --require-token

system.http 不发送自定义请求头,所以在一个新的 version(比如 rules-v2-auth)下,每次调用都会被拒绝:

Gate: BLOCK (exit 3)
│ valid-format   │ format_valid     │ INSUFFICIENT_EVIDENCE │ no_observations │
│ intent_correct   │          │ [0.0%, 100.0%] │ 0 / 0 observed · 20 missing · 0 excluded │

而 oloproof inspect RUN_ID --failures 在每个用例上都显示 execution ERROR: TransientError: system returned HTTP 401。执行错误是缺失的证据,而不是失败:什么都没有被观测到,所以每条规则都是 INSUFFICIENT_EVIDENCE。

变通办法是一个自己发出请求的 Python 可调用对象。client.py 从环境变量中添加请求头,所以令牌永远不会进入 oloproof.yaml 或任何已存储的记录:

URL = os.environ.get("INTAKE_URL", "http://127.0.0.1:8765/extract")


def extract(case: dict[str, Any]) -> dict[str, Any]:
    headers = {"Authorization": f"Bearer {os.environ['INTAKE_TOKEN']}"}
    response = httpx.post(URL, json=case, headers=headers, timeout=30)
    response.raise_for_status()
    return response.json()["result"]

把 oloproof.yaml 中的 http: 块替换为:

system:
  name: order-intake
  version: rules-v2
  callable: client:extract

并在环境中带上令牌运行:

INTAKE_TOKEN=s3cret oloproof run

全部 20 个用例再次被观测到,门禁放行。可调用对象的缓存跟随它自己的源代码和 version,而不是它背后的服务,所以同样的规则适用:部署时修改 version。

故障排除

症状原因与修复
每个用例都出现 system connection failed (ConnectError)服务没有运行,或监听在另一个端口上。
system connection failed (RemoteProtocolError)该端口上有别的东西在应答。选一个空闲端口。
KeyError: "missing output path 'results'"output_path 指定了响应中没有的字段。
system returned HTTP 401 或 403端点需要凭据:使用可调用对象的变通办法。
你部署了变更,但缓存行显示全部命中version 没有变化,所以复用了已存储的输出。
an HTTP system needs a declared version在 system 或 system.http 下添加 version。

局限

  • system.http 上没有自定义请求头、身份验证、查询参数或请求模板:用例的 input 原样作为 JSON 请求体。其他任何需求都请使用可调用对象。
  • 响应必须是 JSON。流式响应不会作为流来读取。
  • Oloproof 无法检测部署;声明的 version 就是应答者的全部身份。
  • 你的服务自己负责其状态:Oloproof 发送请求并记录应答;它不会重置、隔离或回滚请求所改变的任何东西。