指南
配置参考
oloproof.yaml 和 release.yaml 的每个字段,附带其类型、默认值、可接受的取值和示例,均取自读取这些文件的模型。用它来查找某个字段;要学习工作流程,请阅读快速入门和门禁页面。
两个文件都会在任何东西运行之前被校验。未知字段、拼错的字段或类型错误的值都属于配置错误,命令以 2 退出,不执行任何用例。两个文件都有 JSON Schema,能读取 JSON Schema 的编辑器可以用它们进行补全。已安装的软件包会把它们连同结果模式一起写入当前目录下的 schemas/v1/:python -m oloproof_core.models.schema_export(这两个文件是 project_config.schema.json 和 release_policy.schema.json)。
在下面的表格中,“必需”表示缺少该字段时文件会被拒绝;其他字段显示的是省略时所使用的值。
oloproof.yaml 概览
一个小而完整的项目。它在本地运行一个 Python 函数,不需要网络,也不需要密钥,正是 oloproof init 生成的骨架结构。
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label顶层字段
| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| version | 1 | 1 | 文件格式版本。只有 1。 |
| project | 字符串 | 必需 | 项目名称,显示在报告中,并在推送时使用。 |
| dataset | 路径 | 必需 | 测试套件文件,JSONL 格式,相对于项目。它的行在测试套件中有说明。 |
| system | 映射 | 必需 | 被测系统。见下文。 |
| concurrency | 映射 | system: 8、judge: 4 | 同时运行多少个系统调用和评判调用。 |
| evaluators | 列表 | 必需,至少一个 | 在每个用例上度量什么。每个条目都有一个 type。 |
| metrics | 列表 | 空 | 在每个评估器判据本身已构成的指标之外的额外指标。 |
| predictive | 映射 | 无 | 分类器的标签、分数和真实值在哪里。见预测模型。 |
| slices | 字符串列表 | 空 | 探索性切片:metadata.<key>、relevant_position 或 context_truncated。它们从不进入门禁。见切片。 |
| min_slice_support | 整数,至少为 1 | 30 | 符合条件的用例少于这个数时,切片显示其估计值,但没有区间。 |
| replicates | 整数,至少为 1 | 1 | 每个用例度量这么多次。用例仍然是单位:在计算任何区间之前,重复会在用例内部聚合。 |
| pricing | 列表 | 空 | 你按模型为每百万 token 支付的费用。没有它时,成本以 token 报告,从不以美元报告。 |
| egress | 字符串列表 | 空 | oloproof push 可以把哪些原始内容发送到托管工作区。见结果与执行。 |
concurrency
| 字段 | 类型 | 默认值 |
|---|---|---|
| system | 整数,至少为 1 | 8 |
| judge | 整数,至少为 1 | 4 |
pricing 条目
Oloproof 不附带价格表。每个条目指定的模型名称必须与评估器的 model: 完全一致。
| 字段 | 类型 | 默认值 |
|---|---|---|
| model | 字符串 | 必需 |
| input_per_mtok | 数字,0 或以上 | 必需 |
| output_per_mtok | 数字,0 或以上 | 必需 |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
pricing:
- model: my-judge-model
input_per_mtok: 0.15
output_per_mtok: 0.6
egress: [raw_outputs]system
系统需要且只需要 callable、http 或 rag 中的一个。
| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| name | 字符串 | 必需 | 系统的名称。是其版本身份的一部分。 |
| version | 字符串 | 无 | 你为这个版本取的标签。HTTP 系统必需。它是身份的一部分,所以修改它会使缓存的执行失效。 |
| callable | module:attribute | 无 | 一个 Python 函数,同步或异步均可。它接收用例的 input 并返回输出。 |
| http | 映射 | 无 | 对每个用例调用一次的端点。见下文。 |
| rag | 映射 | 无 | 用 @rag_system 声明的分阶段 RAG 类。见下文。 |
| config | 映射 | 空 | 与系统版本一同记录的自由格式设置。修改它们会改变版本。 |
| code_paths | glob 模式列表 | 空 | 其内容计入可调用对象系统版本的源文件。没有它时,只对可调用对象自身的模块计算哈希。 |
| timeout_s | 大于 0 的数字 | 120 | 可调用对象系统每次调用的时限。HTTP 系统改用 http.timeout_s。 |
| records | 产物类型列表 | 空 | 可调用对象系统记录的产物类型,例如 retrieval/v1。在 HTTP 或 RAG 系统上会被拒绝。 |
system.http
| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| url | 字符串 | 必需 | 每个用例发送到哪里。 |
| method | GET、POST 或 PUT | POST | HTTP 方法。 |
| output_path | 点分路径 | 无 | JSON 响应中的哪个字段是输出,例如 result.answer。没有时表示整个响应体。 |
| artifacts | 类型到点分路径的映射 | 空 | 作为产物记录的响应字段,例如 retrieval/v1: debug.retrieval。 |
| version | 字符串 | 无 | 当没有 system.version 时用作系统的版本。HTTP 系统需要两者之一。 |
| timeout_s | 大于 0 的数字 | 30 | 每个请求的时限。 |
# oloproof.yaml
version: 1
project: support-api
dataset: datasets/support.jsonl
system:
name: support-api
version: "2026-10-08"
http:
url: http://localhost:8000/answer
output_path: answer
artifacts:
retrieval/v1: debug.retrieval
evaluators:
- type: hit_rate
k: 5请求和响应的契约,以及超时和 HTTP 错误时会发生什么,见结果与执行。
system.rag
| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| object | module:attribute | 必需 | 用 @rag_system 声明的类,或它的一个实例。 |
| depth | 整数,至少为 1 | 类中的值 | 检索返回多少个段落。 |
| top_k | 整数,至少为 1 | 类中的值 | 其中有多少个进入生成。 |
| token_budget | 整数,至少为 1 | 类中的值 | 上下文的 token 上限。需要类的 count_tokens(passage)。 |
| index_version | 字符串 | 类中的值 | 检索身份的一部分。每当重建索引时修改它。 |
这里给出的设置会覆盖类所声明的设置。分阶段系统自己记录 retrieval/v1、context/v1 和 citations/v1 产物,所以与它同时出现的 records 会被拒绝。见 RAG。
evaluators
每个条目都接受一个 type 以及以下两个通用字段:
| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| criterion | 字符串 | 必需,除非该类型有默认值 | 被度量对象的名称。每个判据都是一个指标,规则的 metric: 指的就是它。 |
| on_execution_error | missing 或 fail | missing | 系统调用失败的用例在这个判据上计为什么。missing 把它作为未观测保留在分母中;fail 把它计为失败。 |
fail 只适用于通过/失败型评估器;在分数评估器上使用它属于配置错误。on_execution_error 是 YAML 字段;SDK 评估器类不接受这样的参数,出错的用例计为缺失。
评估器类型
“读取”列出评估器的结论所依赖的内容,这也是其缓存的评判结果所用的键。“SDK”指明 oloproof.evaluators 中的类。
| YAML type | 读取 | SDK | 是否需要网络或密钥 |
|---|---|---|---|
| exact_match | output、expected | ExactMatch | 否 |
| contains | output、expected | Contains | 否 |
| regex | output | Regex | 否 |
| json_schema | output | JsonSchema | 否 |
| rubric_judge | input、output、expected | RubricJudge | 是,一个模型提供方 |
| model_classifier | output(或 text 指定的字段),可选 premise | 仅限 YAML | 是,一个兼容 TEI 的服务器 |
| probability_judge | 用例和输出 | 仅限 YAML | 是,一个返回对数概率的、兼容 OpenAI 的提供方 |
| cascade | 与它的两个阶段相同 | 仅限 YAML | 是 |
| hit_rate、recall、mrr、ndcg | artifacts.retrieval、expected | HitRate、Recall、MRR、NDCG | 否 |
| citation_validity | artifacts.citations、artifacts.context | CitationValidity | 否 |
| groundedness_judge | input、output、artifacts.context | Groundedness | 是 |
| citation_support_judge | input、output、artifacts.context、artifacts.citations | CitationSupport | 是 |
| agent_max_steps | artifacts.agent_trajectory | AgentMaxSteps | 否 |
| agent_tool_called | artifacts.agent_trajectory | AgentToolCalled | 否 |
| agent_no_tool_loop | artifacts.agent_trajectory | AgentNoToolLoop | 否 |
| agent_tool_sequence | artifacts.agent_trajectory、expected | AgentToolSequence | 否 |
| agent_no_undeclared_tool | artifacts.agent_trajectory、expected | AgentNoUndeclaredTool | 否 |
| agent_constraints_satisfied | artifacts.agent_trajectory | AgentConstraintsSatisfied | 否 |
| agent_route | artifacts.agent_trajectory | AgentRoute | 否 |
| agent_tool_permissions | artifacts.agent_trajectory | AgentToolPermissions | 否 |
| agent_max_handoffs | artifacts.agent_trajectory | AgentMaxHandoffs | 否 |
| predictive_correct | output 和 expected 的标签字段 | PredictiveCorrect | 否 |
| predictive_recall | 同上 | PredictiveRecall | 否 |
| predictive_precision | 同上 | PredictivePrecision | 否 |
| predictive_absolute_error | 同上,数值型 | AbsoluteError | 否 |
| predictive_brier | output 的分数字段,expected 的标签 | Brier | 否 |
| predictive_log_loss | 同上 | LogLoss | 否 |
| predictive_ranking | 同上 | PredictiveRanking | 否 |
| 没有 YAML 类型 | artifacts.conversation | ConversationCompleted(仅限 SDK) | 否 |
| 没有 YAML 类型 | expected、artifacts.conversation | ConversationJudge(仅限 SDK) | 是 |
| 没有 YAML 类型 | 你所声明的内容 | @evaluator 和 CustomEvaluator(仅限 SDK) | 由你决定 |
调用托管模型的评判模型会把用例内容发送给该提供方,并由其计费。密钥从 api_key_env 中指定的环境变量读取;Oloproof 从不把它们存储在这些文件中。
确定性评估器
| 类型 | 字段 | 类型 | 默认值 |
|---|---|---|---|
| exact_match | field | 输出中的点分路径 | 无:整个输出 |
| exact_match | expected_field | expected 中的点分路径 | 无:与 field 相同 |
| exact_match | strip | 布尔值 | true |
| exact_match | casefold | 布尔值 | false |
| contains | field、expected_field | 与 exact_match 相同 | 无 |
| regex | pattern | 正则表达式 | 必需 |
| regex | field | 点分路径 | 无 |
| regex | pass_if | match 或 no_match | match |
| json_schema | schema | 内联的 JSON Schema,或相对于项目的 JSON 文件路径 | 必需 |
| json_schema | field | 点分路径 | 无 |
模型评判器
rubric_judge、groundedness_judge 和 citation_support_judge 共享这些字段。rubric_judge 需要且只需要 rubric_file 或 rubric_text 中的一个;两个 RAG 评判器最多接受一个,否则使用内置的评分细则。它们的 criterion 默认分别为 groundedness 和 citation_support。
| 字段 | 类型 | 默认值 |
|---|---|---|
| provider | anthropic、openai 或 openai_compatible | 必需 |
| model | 字符串 | 必需 |
| rubric_file | 路径 | 无 |
| rubric_text | 字符串 | 无 |
| api_key_env | 环境变量名 | ANTHROPIC_API_KEY 或 OPENAI_API_KEY |
| base_url | URL | 提供方的默认地址 |
| temperature | 数字 | 0 |
| max_tokens | 整数,至少为 1 | 512 |
| timeout_s | 大于 0 的数字 | 60 |
probability_judge 提出一个有类型的问题,并读取模型的概率:
| 字段 | 类型 | 默认值 |
|---|---|---|
| provider | openai 或 openai_compatible | 必需 |
| model | 字符串 | 必需 |
| question | 字符串 | 必需 |
| form | yes_no、choice 或 score | 必需 |
| min_probability | (0, 1] 中的数字 | 必需 |
| options | 回答到描述的映射 | 用于 choice |
| pass_options | 回答列表 | 用于 choice |
| levels | 等级到描述的映射,最低的在前 | 用于 score |
| pass_at_least | 一个等级 | 用于 score |
| calibration | slope(大于 0)、intercept、from_version | 无 |
| api_key_env、base_url | 同上 | 无 |
| timeout_s | 大于 0 的数字 | 60 |
cascade 先运行一个低成本的评判模型,并把不确定的用例升级处理:
| 字段 | 类型 | 默认值 |
|---|---|---|
| first | 一个 probability_judge 条目 | 必需 |
| then | 一个 rubric_judge 或 probability_judge 条目 | 必需 |
| escalate_between | 两个概率 | 必需 |
各阶段评判的是级联自己的 criterion;指定了不同判据的阶段会被拒绝。
model_classifier 用兼容 TEI 的服务器上的训练好的模型为文本打分:
| 字段 | 类型 | 默认值 |
|---|---|---|
| model | 字符串 | 必需 |
| base_url | URL | 必需 |
| label | 要读取的分类器标签 | 必需 |
| min_score 或 max_score | [0, 1] 中的数字,恰好一个 | 必需 |
| text | 对哪个字段进行分类 | output |
| premise | 第二段文本,用于句对分类器 | 无 |
| api_key_env | 环境变量名 | 无 |
| timeout_s | 大于 0 的数字 | 30 |
RAG 评估器
| 类型 | 字段 | 类型 | 默认值 |
|---|---|---|---|
| hit_rate、recall、mrr、ndcg | k | 整数,至少为 1 | hit_rate 和 recall 为 5,mrr 和 ndcg 为 10 |
| hit_rate、recall、mrr、ndcg | relevance_unit | doc 或 chunk | doc |
| hit_rate、recall、mrr、ndcg | criterion | 字符串 | <type>_at_<k>,例如 hit_rate_at_5 |
| citation_validity | require_citations | 布尔值 | false |
| citation_validity | criterion | 字符串 | citations_valid |
智能体评估器
| 类型 | 字段 | 类型 | 默认值 |
|---|---|---|---|
| agent_max_steps | max_steps | 整数,至少为 1 | 必需 |
| agent_tool_called | tool_name | 字符串 | 必需 |
| agent_tool_called | min_calls | 整数,至少为 1 | 1 |
| agent_no_tool_loop | max_repeats | 整数,至少为 1 | 2 |
| agent_tool_sequence | ordered | 布尔值 | true |
| agent_constraints_satisfied | constraints | 约束名称列表 | 空 |
| agent_tool_permissions | permissions | 智能体到允许工具的映射 | 必需 |
| agent_max_handoffs | max_handoffs | 整数,0 或以上 | 必需 |
每种智能体类型都有默认的 criterion,所以可以省略:它自己的类型名称,或由其设置构建的名称(agent_steps_le_8、agent_tool_lookup_called、agent_handoffs_le_2)。见智能体。
预测评估器
| 类型 | 字段 | 类型 | 默认值 |
|---|---|---|---|
| predictive_correct、predictive_recall、predictive_precision | positive | 任意 JSON 值 | true,或 predictive: 块中的值 |
| 同上 | field | 输出字段 | label,或 predictive.label_field |
| 同上 | expected_field | 预期字段 | label,或 predictive.expected_field |
| predictive_absolute_error | target_range | 两个数字 | 必需 |
| predictive_absolute_error | field、expected_field | 同上 | label |
| predictive_brier、predictive_log_loss、predictive_ranking | positive | 任意 JSON 值 | true,或块中的值 |
| 同上 | field | 输出字段 | score,或 predictive.score_field |
| 同上 | expected_field | 预期字段 | label,或块中的值 |
| predictive_log_loss | clip | (0, 0.5) 中的数字 | 必需 |
预测评估器如果没有写出 positive、field 或 expected_field,就从 predictive: 块中取值;它自己写出的值会被保留。
predictive
| 字段 | 类型 | 默认值 |
|---|---|---|
| label_field | 字符串 | label |
| score_field | 字符串 | score |
| expected_field | 字符串 | label |
| positive | 任意 JSON 值 | true |
| calibration_bins | 整数,至少为 1 | 10 |
| thresholds | 数字列表 | 空 |
| average | macro 或 micro | 无:不聚合 |
metrics
每个评估器判据本身已经是一个指标。一个 metrics: 条目再增加一个,以 type 区分。
| type | 字段 | 它是什么 |
|---|---|---|
| quantile | id、source、(0, 1) 中的 quantile | latency_ms、input_tokens、output_tokens、cost_usd、agent_steps 或 agent_tool_calls 的一个分位数。 |
| ranking | id、criterion、statistic:roc_auc 或 average_precision | 基于某个排序判据分数顺序的统计量。 |
| human_score、human_preference | id | 被拒绝:目前还没有已认可的方法读取这些标签。 |
| cost_per_accepted | id、criterion、cost_ceiling_usd、cost_ceiling_source | 在其接线通过审计认可之前被拒绝。 |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
metrics:
- id: latency_p95
type: quantile
source: latency_ms
quantile: 0.95release.yaml
发布策略:哪些规则作出决策,哪些决策会阻止发布。省略的设置保持默认值,所以只指定了规则的策略仍会在 FAIL、INSUFFICIENT_EVIDENCE 和 MANUAL_REVIEW 时阻止。
# release.yaml
version: 1
rules:
- id: label_accuracy
metric: correct_label
min: 0.8| 字段 | 类型 | 默认值 | 它是什么 |
|---|---|---|---|
| version | 1 | 1 | 文件格式版本。 |
| confidence_level | 概率 | 0.95 | 规则读取的每个区间的置信水平。 |
| block_on | 决策状态列表 | FAIL、INSUFFICIENT_EVIDENCE、MANUAL_REVIEW | 让门禁阻止并设定退出码的状态。 |
| warn_on | 决策状态列表 | 空 | 只警告而不阻止的状态。不能与 block_on 重叠。 |
| block_on_partial_run | 布尔值 | true | 没有完成的运行是否以退出码 5 阻止。 |
| require_validated_evaluators | 布尔值 | true | 基于模型评判器的规则是否在评判模型对照人工标签验证之前暂缓决策。确定性评估器不受此限。 |
| minimum_evaluator_agreement | [0, 1] 中的数字 | 无 | 评判模型在获准验证之前,与人工标签的一致性按其下界必须达到的水平。 |
| maximum_evaluator_bias | (0, 1] 中的数字 | 无 | 评判模型在获准验证之前,其通过率与人工通过率之间允许的最大差距。 |
| allow_approximate_methods | 布尔值 | false | 规则是否可以依据引擎标记为近似的区间(聚类二值区间)作出决策。否则它显示 MANUAL_REVIEW。 |
| min_clusters | 整数,至少为 10 | 20 | 聚类少于这个数时,聚类规则显示 INSUFFICIENT_EVIDENCE。 |
| difference_method | bounded_paired_difference@1 或 conditional_exact_paired_difference@1 | 无:使用第一个 | 由哪种已认可的方法为配对二值比率差异设界。 |
| early_stopping | 布尔值 | false | 分批运行用例,一旦每条规则都已决策就停止。见门禁。 |
| early_stopping_seed | 整数,0 或以上 | 无 | 用例顺序的种子。 |
| early_stopping_batch_size | 整数,至少为 1 | 25 | 每批的用例数。 |
| rules | 列表 | 必需,至少一条 | 规则。见下文。 |
| families | 列表 | 空 | 其错误 FAIL 被共同控制的规则。 |
| review_rule | 映射 | 无 | 被拒绝:其接线尚未获得认可。 |
rules
一个列表同时容纳两种规则。运行规则需要且只需要 min、max 或 max_failures 中的一个。比较规则指明它的 kind,并对两次运行之间的差异作出决策;见比较规则。
| 字段 | 类型 | 默认值 | 适用于 |
|---|---|---|---|
| id | 字符串 | 必需 | 全部 |
| metric | 指标 id 或判据 | 必需 | 全部 |
| kind | interval_threshold、observed_count、superiority、non_inferiority、equivalence | 运行规则可推断 | 全部 |
| min | 数字 | 无 | 运行规则:当区间下界至少为该值时 PASS |
| max | 数字 | 无 | 运行规则:当区间上界至多为该值时 PASS |
| max_failures | 整数,0 或以上 | 无 | observed_count:基于已执行测试套件的计数,没有区间 |
| margin | 大于 0 的数字,以指标的单位计 | 无 | non_inferiority 和 equivalence;在 superiority 上被拒绝 |
| direction | min 或 max | min | 仅 non_inferiority:越高越好还是越低越好 |
| max_missing_fraction | [0, 1] 中的数字 | 无 | 区间规则和比较规则 |
| requires_manual_review | 布尔值 | false | 全部:该规则总是显示 MANUAL_REVIEW |
| scope | global 或一个切片 | global | 区间规则和比较规则 |
| min_support | 整数,至少为 1 | 无 | 针对切片的比较规则 |
families
| 字段 | 类型 | 默认值 |
|---|---|---|
| id | 字符串 | 必需 |
| correction | holm | holm |
| rules | 规则 id 列表 | 必需,至少一个 |
# release.yaml
version: 1
warn_on: [INSUFFICIENT_EVIDENCE]
block_on: [FAIL, MANUAL_REVIEW]
rules:
- id: label_accuracy
metric: correct_label
min: 0.8
max_missing_fraction: 0.05
- id: no_regression
metric: correct_label
kind: non_inferiority
margin: 0.02产物类型
产物是系统在其输出旁边写下的有类型记录,例如它检索到了什么。类型是一个小写名称,可选带有版本,匹配 ^[a-z][a-z0-9_]*(/v[1-9][0-9]*)?$。需要某个产物的评估器会指明它,而如果系统没有声明某个必需的类型,运行会在开始之前被拒绝,而不是把每个用例都计为缺失。
| 类型 | 由谁写入 | 由谁需要 |
|---|---|---|
| retrieval/v1 | current_case().retrieval(...)、一个 @rag_system 或 http.artifacts | hit_rate、recall、mrr、ndcg |
| context/v1 | current_case().context(...) 或一个 @rag_system | citation_validity、groundedness_judge、citation_support_judge |
| citations/v1 | current_case().citations(...) 或一个 @rag_system | citation_validity、citation_support_judge |
| agent_trajectory/v1 | current_case().agent_trajectory(...) | 每个 agent_* 评估器,以及 agent_steps 和 agent_tool_calls 来源 |
| conversation/v1 | current_case().artifact(CONVERSATION, ...) | ConversationCompleted、ConversationJudge |
| stage_timings/v1 | 一个 @rag_system | 无;显示在延迟旁边 |
可调用对象系统在 records:(或 @system(records=...))中声明它记录的类型;HTTP 系统在 http.artifacts 中声明;分阶段 RAG 系统自己记录。
版本、缓存键与失效
Oloproof 复用输入没有变化的工作,并根据内容摘要来判定什么是“没有变化”。每个摘要都由引擎计算,并随运行一起记录。
| 记录 | 当以下内容完全相同时被复用 |
|---|---|
| 系统版本 | name、version、config,以及一个代码摘要:可调用对象的模块源代码(或 code_paths 匹配到的每个文件),HTTP 系统的 url、method、output_path 和 artifacts |
| 执行 | 系统版本、用例的 input 和重复序号。只有成功的执行会被复用。 |
| 评判结果 | 评估器版本(它的类型和每一项设置),以及它读取的每个字段的摘要,如评估器表中所列 |
| 分析 | 分析计划、指标、置信水平、测试套件摘要以及它计入的每个输入 |
| 门禁 | 每个分析、策略摘要、运行是否完成,以及决策所引用的每个评估器的有效状态 |
Oloproof 看不到的东西需要由你来声明:
- HTTP 系统的行为在服务器上。每当 URL 背后的东西发生变化时,修改 system.version,否则旧的缓存输出会代表新的系统。
- 可调用对象的辅助模块只有在被 code_paths 匹配时才会计算哈希。没有它时,编辑辅助模块不会改变版本。
- 方法或可调用对象必须声明版本,并且当对象的状态改变时,版本也必须改变。
- RAG 索引由 index_version 标识;重建索引时修改它。
- 模型评判器的身份是它的设置,而不是提供方的权重。提供方在同一名称背后更新模型,缓存是察觉不到的。
- 自定义 @evaluator 对定义它的模块文件计算哈希,并且只有在它声明 cacheable=True 时,它的评判结果才会在多次运行之间被复用。内置的评分细则评判器是可缓存的;确定性评估器会被重新计算,这成本很低。
缓存的工作保存在项目的本地存储中,即 oloproof.yaml 旁边的 .oloproof/store.sqlite(或 OLOPROOF_HOME 之下)。删除存储会丢弃所有缓存和所有运行。在托管工作区中,引擎不复用缓存的执行、评判结果或分析,因为推送可以写入它们;它会重新计算。