指南
設定參考
oloproof.yaml 和 release.yaml 的每個欄位,附帶其類型、預設值、可接受的取值和範例,均取自讀取這些檔案的模型。用它來查找某個欄位;要學習工作流程,請閱讀快速入門和閘門頁面。
兩個檔案都會在任何東西執行之前被校驗。未知欄位、拼錯的欄位或類型錯誤的值都屬於設定錯誤,命令以 2 結束,不執行任何案例。兩個檔案都有 JSON Schema,能讀取 JSON Schema 的編輯器可以用它們進行補全。已安裝的軟體包會把它們連同結果模式一起寫入當前目錄下的 schemas/v1/:python -m oloproof_core.models.schema_export(這兩個檔案是 project_config.schema.json 和 release_policy.schema.json)。
在下面的表格中,“必需”表示缺少該欄位時檔案會被拒絕;其他欄位顯示的是省略時所使用的值。
oloproof.yaml 概覽
一個小而完整的專案。它在本機執行一個 Python 函式,不需要網路,也不需要密鑰,正是 oloproof init 生成的骨架結構。
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label頂層欄位
| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| version | 1 | 1 | 檔案格式版本。只有 1。 |
| project | 字串 | 必需 | 專案名稱,顯示在報告中,並在推送時使用。 |
| dataset | 路徑 | 必需 | 套件檔案,JSONL 格式,相對於專案。它的行在套件中有說明。 |
| system | 映射 | 必需 | 被測系統。見下文。 |
| concurrency | 映射 | system: 8、judge: 4 | 同時執行多少個系統呼叫和評判呼叫。 |
| evaluators | 清單 | 必需,至少一個 | 在每個案例上量測什麼。每個條目都有一個 type。 |
| metrics | 清單 | 空 | 在每個評估器判據本身已構成的指標之外的額外指標。 |
| predictive | 映射 | 無 | 分類器的標籤、分數和真實值在哪裡。見預測模型。 |
| slices | 字串清單 | 空 | 探索性切片:metadata.<key>、relevant_position 或 context_truncated。它們從不進入閘門。見切片。 |
| min_slice_support | 整數,至少為 1 | 30 | 符合條件的案例少於這個數時,切片顯示其估計值,但沒有區間。 |
| replicates | 整數,至少為 1 | 1 | 每個案例量測這麼多次。案例仍然是單位:在計算任何區間之前,重複會在案例內部聚合。 |
| pricing | 清單 | 空 | 你按模型為每百萬 token 支付的費用。沒有它時,成本以 token 報告,從不以美元報告。 |
| egress | 字串清單 | 空 | oloproof push 可以把哪些原始內容發送到託管工作區。見結果與執行。 |
concurrency
| 欄位 | 類型 | 預設值 |
|---|---|---|
| system | 整數,至少為 1 | 8 |
| judge | 整數,至少為 1 | 4 |
pricing 條目
Oloproof 不附帶價格表。每個條目指定的模型名稱必須與評估器的 model: 完全一致。
| 欄位 | 類型 | 預設值 |
|---|---|---|
| model | 字串 | 必需 |
| input_per_mtok | 數字,0 或以上 | 必需 |
| output_per_mtok | 數字,0 或以上 | 必需 |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
pricing:
- model: my-judge-model
input_per_mtok: 0.15
output_per_mtok: 0.6
egress: [raw_outputs]system
系統需要且只需要 callable、http 或 rag 中的一個。
| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| name | 字串 | 必需 | 系統的名稱。是其版本身份的一部分。 |
| version | 字串 | 無 | 你為這個版本取的標籤。HTTP 系統必需。它是身份的一部分,所以修改它會使快取的執行失效。 |
| callable | module:attribute | 無 | 一個 Python 函式,同步或非同步均可。它接收案例的 input 並返回輸出。 |
| http | 映射 | 無 | 對每個案例呼叫一次的端點。見下文。 |
| rag | 映射 | 無 | 用 @rag_system 宣告的分階段 RAG 類。見下文。 |
| config | 映射 | 空 | 與系統版本一同記錄的自由格式設置。修改它們會改變版本。 |
| code_paths | glob 模式清單 | 空 | 其內容計入可呼叫物件系統版本的源檔案。沒有它時,只對可呼叫物件自身的模組計算雜湊。 |
| timeout_s | 大於 0 的數字 | 120 | 可呼叫物件系統每次呼叫的時限。HTTP 系統改用 http.timeout_s。 |
| records | 產物類型清單 | 空 | 可呼叫物件系統記錄的產物類型,例如 retrieval/v1。在 HTTP 或 RAG 系統上會被拒絕。 |
system.http
| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| url | 字串 | 必需 | 每個案例發送到哪裡。 |
| method | GET、POST 或 PUT | POST | HTTP 方法。 |
| output_path | 點分路徑 | 無 | JSON 響應中的哪個欄位是輸出,例如 result.answer。沒有時表示整個響應體。 |
| artifacts | 類型到點分路徑的映射 | 空 | 作為產物記錄的響應欄位,例如 retrieval/v1: debug.retrieval。 |
| version | 字串 | 無 | 當沒有 system.version 時用作系統的版本。HTTP 系統需要兩者之一。 |
| timeout_s | 大於 0 的數字 | 30 | 每個請求的時限。 |
# oloproof.yaml
version: 1
project: support-api
dataset: datasets/support.jsonl
system:
name: support-api
version: "2026-10-08"
http:
url: http://localhost:8000/answer
output_path: answer
artifacts:
retrieval/v1: debug.retrieval
evaluators:
- type: hit_rate
k: 5請求和響應的契約,以及超時和 HTTP 錯誤時會發生什麼,見結果與執行。
system.rag
| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| object | module:attribute | 必需 | 用 @rag_system 宣告的類,或它的一個實例。 |
| depth | 整數,至少為 1 | 類中的值 | 檢索返回多少個段落。 |
| top_k | 整數,至少為 1 | 類中的值 | 其中有多少個進入生成。 |
| token_budget | 整數,至少為 1 | 類中的值 | 上下文的 token 上限。需要類的 count_tokens(passage)。 |
| index_version | 字串 | 類中的值 | 檢索身份的一部分。每當重建索引時修改它。 |
這裡給出的設置會覆蓋類所宣告的設置。分階段系統自己記錄 retrieval/v1、context/v1 和 citations/v1 產物,所以與它同時出現的 records 會被拒絕。見 RAG。
evaluators
每個條目都接受一個 type 以及以下兩個通用欄位:
| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| criterion | 字串 | 必需,除非該類型有預設值 | 被量測物件的名稱。每個判據都是一個指標,規則的 metric: 指的就是它。 |
| on_execution_error | missing 或 fail | missing | 系統呼叫失敗的案例在這個判據上計為什麼。missing 把它作為未觀測保留在分母中;fail 把它計為失敗。 |
fail 只適用於通過/失敗型評估器;在分數評估器上使用它屬於設定錯誤。on_execution_error 是 YAML 欄位;SDK 評估器類不接受這樣的參數,出錯的案例計為缺失。
評估器類型
“讀取”列出評估器的結論所依賴的內容,這也是其快取的評判結果所用的鍵。“SDK”指明 oloproof.evaluators 中的類。
| YAML type | 讀取 | SDK | 是否需要網路或密鑰 |
|---|---|---|---|
| exact_match | output、expected | ExactMatch | 否 |
| contains | output、expected | Contains | 否 |
| regex | output | Regex | 否 |
| json_schema | output | JsonSchema | 否 |
| rubric_judge | input、output、expected | RubricJudge | 是,一個模型提供者 |
| model_classifier | output(或 text 指定的欄位),可選 premise | 僅限 YAML | 是,一個兼容 TEI 的伺服器 |
| probability_judge | 案例和輸出 | 僅限 YAML | 是,一個返回對數機率的、兼容 OpenAI 的提供者 |
| cascade | 與它的兩個階段相同 | 僅限 YAML | 是 |
| hit_rate、recall、mrr、ndcg | artifacts.retrieval、expected | HitRate、Recall、MRR、NDCG | 否 |
| citation_validity | artifacts.citations、artifacts.context | CitationValidity | 否 |
| groundedness_judge | input、output、artifacts.context | Groundedness | 是 |
| citation_support_judge | input、output、artifacts.context、artifacts.citations | CitationSupport | 是 |
| agent_max_steps | artifacts.agent_trajectory | AgentMaxSteps | 否 |
| agent_tool_called | artifacts.agent_trajectory | AgentToolCalled | 否 |
| agent_no_tool_loop | artifacts.agent_trajectory | AgentNoToolLoop | 否 |
| agent_tool_sequence | artifacts.agent_trajectory、expected | AgentToolSequence | 否 |
| agent_no_undeclared_tool | artifacts.agent_trajectory、expected | AgentNoUndeclaredTool | 否 |
| agent_constraints_satisfied | artifacts.agent_trajectory | AgentConstraintsSatisfied | 否 |
| agent_route | artifacts.agent_trajectory | AgentRoute | 否 |
| agent_tool_permissions | artifacts.agent_trajectory | AgentToolPermissions | 否 |
| agent_max_handoffs | artifacts.agent_trajectory | AgentMaxHandoffs | 否 |
| predictive_correct | output 和 expected 的標籤欄位 | PredictiveCorrect | 否 |
| predictive_recall | 同上 | PredictiveRecall | 否 |
| predictive_precision | 同上 | PredictivePrecision | 否 |
| predictive_absolute_error | 同上,數值型 | AbsoluteError | 否 |
| predictive_brier | output 的分數欄位,expected 的標籤 | Brier | 否 |
| predictive_log_loss | 同上 | LogLoss | 否 |
| predictive_ranking | 同上 | PredictiveRanking | 否 |
| 沒有 YAML 類型 | artifacts.conversation | ConversationCompleted(僅限 SDK) | 否 |
| 沒有 YAML 類型 | expected、artifacts.conversation | ConversationJudge(僅限 SDK) | 是 |
| 沒有 YAML 類型 | 你所宣告的內容 | @evaluator 和 CustomEvaluator(僅限 SDK) | 由你決定 |
呼叫託管模型的評審會把案例內容發送給該提供者,並由其計費。密鑰從 api_key_env 中指定的環境變數讀取;Oloproof 從不把它們儲存在這些檔案中。
確定性評估器
| 類型 | 欄位 | 類型 | 預設值 |
|---|---|---|---|
| exact_match | field | 輸出中的點分路徑 | 無:整個輸出 |
| exact_match | expected_field | expected 中的點分路徑 | 無:與 field 相同 |
| exact_match | strip | 布爾值 | true |
| exact_match | casefold | 布爾值 | false |
| contains | field、expected_field | 與 exact_match 相同 | 無 |
| regex | pattern | 正則表達式 | 必需 |
| regex | field | 點分路徑 | 無 |
| regex | pass_if | match 或 no_match | match |
| json_schema | schema | 內聯的 JSON Schema,或相對於專案的 JSON 檔案路徑 | 必需 |
| json_schema | field | 點分路徑 | 無 |
模型評審
rubric_judge、groundedness_judge 和 citation_support_judge 共享這些欄位。rubric_judge 需要且只需要 rubric_file 或 rubric_text 中的一個;兩個 RAG 評審最多接受一個,否則使用內置的評分準則。它們的 criterion 預設分別為 groundedness 和 citation_support。
| 欄位 | 類型 | 預設值 |
|---|---|---|
| provider | anthropic、openai 或 openai_compatible | 必需 |
| model | 字串 | 必需 |
| rubric_file | 路徑 | 無 |
| rubric_text | 字串 | 無 |
| api_key_env | 環境變數名 | ANTHROPIC_API_KEY 或 OPENAI_API_KEY |
| base_url | URL | 提供者的預設地址 |
| temperature | 數字 | 0 |
| max_tokens | 整數,至少為 1 | 512 |
| timeout_s | 大於 0 的數字 | 60 |
probability_judge 提出一個有類型的問題,並讀取模型的機率:
| 欄位 | 類型 | 預設值 |
|---|---|---|
| provider | openai 或 openai_compatible | 必需 |
| model | 字串 | 必需 |
| question | 字串 | 必需 |
| form | yes_no、choice 或 score | 必需 |
| min_probability | (0, 1] 中的數字 | 必需 |
| options | 回答到描述的映射 | 用於 choice |
| pass_options | 回答清單 | 用於 choice |
| levels | 等級到描述的映射,最低的在前 | 用於 score |
| pass_at_least | 一個等級 | 用於 score |
| calibration | slope(大於 0)、intercept、from_version | 無 |
| api_key_env、base_url | 同上 | 無 |
| timeout_s | 大於 0 的數字 | 60 |
cascade 先執行一個低成本的評審,並把不確定的案例升級處理:
| 欄位 | 類型 | 預設值 |
|---|---|---|
| first | 一個 probability_judge 條目 | 必需 |
| then | 一個 rubric_judge 或 probability_judge 條目 | 必需 |
| escalate_between | 兩個機率 | 必需 |
各階段評判的是級聯自己的 criterion;指定了不同判據的階段會被拒絕。
model_classifier 用兼容 TEI 的伺服器上的訓練好的模型為文字打分:
| 欄位 | 類型 | 預設值 |
|---|---|---|
| model | 字串 | 必需 |
| base_url | URL | 必需 |
| label | 要讀取的分類器標籤 | 必需 |
| min_score 或 max_score | [0, 1] 中的數字,恰好一個 | 必需 |
| text | 對哪個欄位進行分類 | output |
| premise | 第二段文字,用於句對分類器 | 無 |
| api_key_env | 環境變數名 | 無 |
| timeout_s | 大於 0 的數字 | 30 |
RAG 評估器
| 類型 | 欄位 | 類型 | 預設值 |
|---|---|---|---|
| hit_rate、recall、mrr、ndcg | k | 整數,至少為 1 | hit_rate 和 recall 為 5,mrr 和 ndcg 為 10 |
| hit_rate、recall、mrr、ndcg | relevance_unit | doc 或 chunk | doc |
| hit_rate、recall、mrr、ndcg | criterion | 字串 | <type>_at_<k>,例如 hit_rate_at_5 |
| citation_validity | require_citations | 布爾值 | false |
| citation_validity | criterion | 字串 | citations_valid |
代理評估器
| 類型 | 欄位 | 類型 | 預設值 |
|---|---|---|---|
| agent_max_steps | max_steps | 整數,至少為 1 | 必需 |
| agent_tool_called | tool_name | 字串 | 必需 |
| agent_tool_called | min_calls | 整數,至少為 1 | 1 |
| agent_no_tool_loop | max_repeats | 整數,至少為 1 | 2 |
| agent_tool_sequence | ordered | 布爾值 | true |
| agent_constraints_satisfied | constraints | 約束名稱清單 | 空 |
| agent_tool_permissions | permissions | 代理到允許工具的映射 | 必需 |
| agent_max_handoffs | max_handoffs | 整數,0 或以上 | 必需 |
每種代理類型都有預設的 criterion,所以可以省略:它自己的類型名稱,或由其設置建置的名稱(agent_steps_le_8、agent_tool_lookup_called、agent_handoffs_le_2)。見代理。
預測評估器
| 類型 | 欄位 | 類型 | 預設值 |
|---|---|---|---|
| predictive_correct、predictive_recall、predictive_precision | positive | 任意 JSON 值 | true,或 predictive: 塊中的值 |
| 同上 | field | 輸出欄位 | label,或 predictive.label_field |
| 同上 | expected_field | 預期欄位 | label,或 predictive.expected_field |
| predictive_absolute_error | target_range | 兩個數字 | 必需 |
| predictive_absolute_error | field、expected_field | 同上 | label |
| predictive_brier、predictive_log_loss、predictive_ranking | positive | 任意 JSON 值 | true,或塊中的值 |
| 同上 | field | 輸出欄位 | score,或 predictive.score_field |
| 同上 | expected_field | 預期欄位 | label,或塊中的值 |
| predictive_log_loss | clip | (0, 0.5) 中的數字 | 必需 |
預測評估器如果沒有寫出 positive、field 或 expected_field,就從 predictive: 塊中取值;它自己寫出的值會被保留。
predictive
| 欄位 | 類型 | 預設值 |
|---|---|---|
| label_field | 字串 | label |
| score_field | 字串 | score |
| expected_field | 字串 | label |
| positive | 任意 JSON 值 | true |
| calibration_bins | 整數,至少為 1 | 10 |
| thresholds | 數字清單 | 空 |
| average | macro 或 micro | 無:不聚合 |
metrics
每個評估器判據本身已經是一個指標。一個 metrics: 條目再增加一個,以 type 區分。
| type | 欄位 | 它是什麼 |
|---|---|---|
| quantile | id、source、(0, 1) 中的 quantile | latency_ms、input_tokens、output_tokens、cost_usd、agent_steps 或 agent_tool_calls 的一個分位數。 |
| ranking | id、criterion、statistic:roc_auc 或 average_precision | 基於某個排序判據分數順序的統計量。 |
| human_score、human_preference | id | 被拒絕:目前還沒有已認可的方法讀取這些標籤。 |
| cost_per_accepted | id、criterion、cost_ceiling_usd、cost_ceiling_source | 在其接線通過審計認可之前被拒絕。 |
# oloproof.yaml
version: 1
project: support-bot
dataset: datasets/support.jsonl
system:
name: support-bot
callable: app.bot:answer
evaluators:
- type: exact_match
criterion: correct_label
field: label
metrics:
- id: latency_p95
type: quantile
source: latency_ms
quantile: 0.95release.yaml
發布政策:哪些規則作出決策,哪些決策會阻止發布。省略的設置保持預設值,所以只指定了規則的政策仍會在 FAIL、INSUFFICIENT_EVIDENCE 和 MANUAL_REVIEW 時阻止。
# release.yaml
version: 1
rules:
- id: label_accuracy
metric: correct_label
min: 0.8| 欄位 | 類型 | 預設值 | 它是什麼 |
|---|---|---|---|
| version | 1 | 1 | 檔案格式版本。 |
| confidence_level | 機率 | 0.95 | 規則讀取的每個區間的信賴水準。 |
| block_on | 決策狀態清單 | FAIL、INSUFFICIENT_EVIDENCE、MANUAL_REVIEW | 讓閘門阻止並設定結束代碼的狀態。 |
| warn_on | 決策狀態清單 | 空 | 只警告而不阻止的狀態。不能與 block_on 重疊。 |
| block_on_partial_run | 布爾值 | true | 沒有完成的執行是否以結束代碼 5 阻止。 |
| require_validated_evaluators | 布爾值 | true | 基於模型評審的規則是否在評審對照人工標籤驗證之前暫緩決策。確定性評估器不受此限。 |
| minimum_evaluator_agreement | [0, 1] 中的數字 | 無 | 評審在獲准驗證之前,與人工標籤的一致性按其下界必須達到的水平。 |
| maximum_evaluator_bias | (0, 1] 中的數字 | 無 | 評審在獲准驗證之前,其通過率與人工通過率之間允許的最大差距。 |
| allow_approximate_methods | 布爾值 | false | 規則是否可以依據引擎標記為近似的區間(叢集二值區間)作出決策。否則它顯示 MANUAL_REVIEW。 |
| min_clusters | 整數,至少為 10 | 20 | 叢集少於這個數時,叢集規則顯示 INSUFFICIENT_EVIDENCE。 |
| difference_method | bounded_paired_difference@1 或 conditional_exact_paired_difference@1 | 無:使用第一個 | 由哪種已認可的方法為配對二值比率差異設界。 |
| early_stopping | 布爾值 | false | 分批執行案例,一旦每條規則都已決策就停止。見閘門。 |
| early_stopping_seed | 整數,0 或以上 | 無 | 案例順序的種子。 |
| early_stopping_batch_size | 整數,至少為 1 | 25 | 每批的案例數。 |
| rules | 清單 | 必需,至少一條 | 規則。見下文。 |
| families | 清單 | 空 | 其錯誤 FAIL 被共同控制的規則。 |
| review_rule | 映射 | 無 | 被拒絕:其接線尚未獲得認可。 |
rules
一個清單同時容納兩種規則。執行規則需要且只需要 min、max 或 max_failures 中的一個。比較規則指明它的 kind,並對兩次執行之間的差異作出決策;見比較規則。
| 欄位 | 類型 | 預設值 | 適用於 |
|---|---|---|---|
| id | 字串 | 必需 | 全部 |
| metric | 指標 id 或判據 | 必需 | 全部 |
| kind | interval_threshold、observed_count、superiority、non_inferiority、equivalence | 執行規則可推斷 | 全部 |
| min | 數字 | 無 | 執行規則:當區間下界至少為該值時 PASS |
| max | 數字 | 無 | 執行規則:當區間上界至多為該值時 PASS |
| max_failures | 整數,0 或以上 | 無 | observed_count:基於已執行套件的計數,沒有區間 |
| margin | 大於 0 的數字,以指標的單位計 | 無 | non_inferiority 和 equivalence;在 superiority 上被拒絕 |
| direction | min 或 max | min | 僅 non_inferiority:越高越好還是越低越好 |
| max_missing_fraction | [0, 1] 中的數字 | 無 | 區間規則和比較規則 |
| requires_manual_review | 布爾值 | false | 全部:該規則總是顯示 MANUAL_REVIEW |
| scope | global 或一個切片 | global | 區間規則和比較規則 |
| min_support | 整數,至少為 1 | 無 | 針對切片的比較規則 |
families
| 欄位 | 類型 | 預設值 |
|---|---|---|
| id | 字串 | 必需 |
| correction | holm | holm |
| rules | 規則 id 清單 | 必需,至少一個 |
# release.yaml
version: 1
warn_on: [INSUFFICIENT_EVIDENCE]
block_on: [FAIL, MANUAL_REVIEW]
rules:
- id: label_accuracy
metric: correct_label
min: 0.8
max_missing_fraction: 0.05
- id: no_regression
metric: correct_label
kind: non_inferiority
margin: 0.02產物類型
產物是系統在其輸出旁邊寫下的有類型記錄,例如它檢索到了什麼。類型是一個小寫名稱,可選帶有版本,匹配 ^[a-z][a-z0-9_]*(/v[1-9][0-9]*)?$。需要某個產物的評估器會指明它,而如果系統沒有宣告某個必需的類型,執行會在開始之前被拒絕,而不是把每個案例都計為缺失。
| 類型 | 由誰寫入 | 由誰需要 |
|---|---|---|
| retrieval/v1 | current_case().retrieval(...)、一個 @rag_system 或 http.artifacts | hit_rate、recall、mrr、ndcg |
| context/v1 | current_case().context(...) 或一個 @rag_system | citation_validity、groundedness_judge、citation_support_judge |
| citations/v1 | current_case().citations(...) 或一個 @rag_system | citation_validity、citation_support_judge |
| agent_trajectory/v1 | current_case().agent_trajectory(...) | 每個 agent_* 評估器,以及 agent_steps 和 agent_tool_calls 來源 |
| conversation/v1 | current_case().artifact(CONVERSATION, ...) | ConversationCompleted、ConversationJudge |
| stage_timings/v1 | 一個 @rag_system | 無;顯示在延遲旁邊 |
可呼叫物件系統在 records:(或 @system(records=...))中宣告它記錄的類型;HTTP 系統在 http.artifacts 中宣告;分階段 RAG 系統自己記錄。
版本、快取鍵與失效
Oloproof 複用輸入沒有變化的工作,並根據內容摘要來判定什麼是“沒有變化”。每個摘要都由引擎計算,並隨執行一起記錄。
| 記錄 | 當以下內容完全相同時被複用 |
|---|---|
| 系統版本 | name、version、config,以及一個代碼摘要:可呼叫物件的模組原始碼(或 code_paths 匹配到的每個檔案),HTTP 系統的 url、method、output_path 和 artifacts |
| 執行 | 系統版本、案例的 input 和重複序號。只有成功的執行會被複用。 |
| 評判結果 | 評估器版本(它的類型和每一項設置),以及它讀取的每個欄位的摘要,如評估器表中所列 |
| 分析 | 分析計劃、指標、信賴水準、套件摘要以及它計入的每個輸入 |
| 閘門 | 每個分析、政策摘要、執行是否完成,以及決策所引用的每個評估器的有效狀態 |
Oloproof 看不到的東西需要由你來宣告:
- HTTP 系統的行為在伺服器上。每當 URL 背後的東西發生變化時,修改 system.version,否則舊的快取輸出會代表新的系統。
- 可呼叫物件的輔助模組只有在被 code_paths 匹配時才會計算雜湊。沒有它時,編輯輔助模組不會改變版本。
- 方法或可呼叫物件必須宣告版本,並且當物件的狀態改變時,版本也必須改變。
- RAG 索引由 index_version 標識;重建索引時修改它。
- 模型評審的身份是它的設置,而不是提供者的權重。提供者在同一名稱背後更新模型,快取是察覺不到的。
- 自定義 @evaluator 對定義它的模組檔案計算雜湊,並且只有在它宣告 cacheable=True 時,它的評判結果才會在多次執行之間被複用。內置的評分準則評審是可快取的;確定性評估器會被重新計算,這成本很低。
快取的工作保存在專案的本機儲存中,即 oloproof.yaml 旁邊的 .oloproof/store.sqlite(或 OLOPROOF_HOME 之下)。刪除儲存會丟棄所有快取和所有執行。在託管工作區中,引擎不復用快取的執行、評判結果或分析,因為推送可以寫入它們;它會重新計算。