Guides
Tutorial: a multi-turn conversation
A runnable walkthrough for evaluating a conversational assistant with the Python SDK: your application replays a scripted conversation against a fresh session, records what it answered as a conversation/v1 artifact, and two evaluators judge each whole conversation. It runs offline with a scripted stand-in for the judge model, then compares a candidate fix, and ends with the alternative of evaluating turns as clustered cases.
Terms such as interval, decision state and release action are defined in Concepts; Recording what a system did covers artifacts.
What Oloproof does and does not do here
| Oloproof does | Your application does |
|---|---|
| store the scripted conversations as the dataset, identical for every system | drive the conversation: ask each scripted turn in order |
| check that every scripted turn was answered (ConversationCompleted) | own the session state, start a fresh session per case, and reset it |
| judge the whole transcript with a model (ConversationJudge) | decide what happens when it cannot continue, and record that it stopped |
| compute intervals, compare two systems, and decide against a policy | record the conversation/v1 artifact |
Oloproof has no user simulator: it never writes a user turn, so the user's side is whatever the dataset scripts. It has no turn-level metric within a recorded conversation, and it cannot replay a recorded conversation against a new system. The conversation evaluators exist in the Python SDK only: ConversationCompleted and ConversationJudge are not evaluator types in oloproof.yaml, so this tutorial uses a script rather than oloproof run.
Prerequisites
- Python 3.11 or later and pip install oloproof, as in the quickstart.
- The example files, which ship with the package. Copy them into a new directory so the run's store lands there:
oloproof init --example conversation ~/oloproof-conversation
cd ~/oloproof-conversation| File | What it is |
|---|---|
| assistant.py | the application under test: a plan assistant with session state |
| systems.py | the adapter: replays a script, records conversation/v1 |
| judge_offline.py | the scripted stand-in for the judge model |
| evaluate.py | runs the evaluation, the comparison and the turns alternative |
| release.yaml | the policy for one run |
| comparison.yaml | the policy for the candidate against the baseline |
| turns_release.yaml | the policy for the turns alternative |
| data/conversations.jsonl | 40 scripted conversations |
| data/turns.jsonl | the same conversations, one case per turn |
No key, no network and no provider cost, until the optional live step at the end.
The application
assistant.py answers questions about three pricing plans. It keeps one piece of state, the plan the conversation is about, so a follow-up such as "Does that include SSO?" can resolve "that". The baseline has a deliberate flaw: it does not remember the plan, so a follow-up is answered about the default plan. A user asking for a person ends the conversation with HandoffRequested.
class PlanAssistant:
def __init__(self, *, remembers_plan):
self.remembers_plan = remembers_plan
self.reset()
def reset(self):
"""Forget everything, so one conversation never leaks into the next."""
self.current_plan = None
def ask(self, question): ...This is the part you replace with your own application: a chatbot client, an agent session, an HTTP session to your service. Whatever it is, it owns its state and its reset; Oloproof sees only what the adapter records.
The dataset: the script is the input
One line of data/conversations.jsonl is one conversation:
{"expected": {"plan": "enterprise"}, "id": "conv_00", "input": {"turns": ["What does the enterprise plan cost?", "Does that include SSO?"]}, "metadata": {"pattern": "pronoun_followup"}}The user's turns are dataset content, covered by the suite's digest and identical for every system measured against them; that is what makes two systems comparable. expected is the reference the judge is shown. The 40 conversations are 24 with a follow-up that names no plan, 12 that name the plan every turn, and 4 that ask for a person at turn two of three.
The adapter
systems.py starts a fresh session per case, asks each scripted turn in order, and records what came back:
from oloproof import CONVERSATION, current_case, system
def replay(case, *, remembers_plan):
script = [str(turn) for turn in case["turns"]]
session = PlanAssistant(remembers_plan=remembers_plan) # a new session per case
turns = []
truncated = False
for index, question in enumerate(script, start=1):
try:
reply = session.ask(question)
except HandoffRequested:
truncated = True # the recording stops here and says so
break
turns.append({"index": index, "asked": question, "answer": reply["answer"]})
current_case().artifact(
CONVERSATION,
{"turns": turns, "declared_turns": len(script), "truncated": truncated},
)
last = turns[-1]["answer"] if turns else None
return {"answer": last, "turns_answered": len(turns)}
@system(name="plan-assistant", version="baseline", records=(CONVERSATION,))
def baseline(case):
return replay(case, remembers_plan=False)
@system(name="plan-assistant", version="candidate-remembers-plan", records=(CONVERSATION,))
def candidate(case):
return replay(case, remembers_plan=True)A fresh session per case matters: Oloproof runs cases concurrently and in no fixed order, and a session shared between cases would let one conversation's state leak into another. records= declares that the system records the artifact; without it the conversation evaluators are refused before anything runs, rather than counting every case as missing.
The conversation/v1 artifact
What the baseline recorded for conv_00, from oloproof export RUN_ID (the bundle's cases.jsonl):
{"declared_turns": 2, "truncated": false, "turns": [{"answer": "The enterprise plan costs a price agreed per contract.", "asked": "What does the enterprise plan cost?", "index": 1, ...}, {"answer": "The starter plan does not include SSO.", "asked": "Does that include SSO?", "index": 2, ...}]}And for a conversation that asked for a person:
{"declared_turns": 3, "truncated": true, "turns": [{"answer": "The team plan costs $20 a month.", "asked": "What does the team plan cost?", "index": 1, ...}]}| Field | Meaning |
|---|---|
| turns[].index | which scripted turn this answers; contiguous from 1 |
| turns[].answer | what the assistant returned, any JSON |
| turns[].asked | optional, for reading only; the engine matches on index |
| turns[].retrieval | optional, what that turn retrieved, in the retrieval/v1 shape |
| declared_turns | how many turns the script declared |
| truncated | the recording stops short of the script, whatever cut it short |
The artifact is checked when it is recorded: a recording with fewer turns than declared must say truncated: true, indexes must be contiguous, and a recording cannot answer more turns than it was asked. A malformed artifact stops the run with a SystemContractError.
The two evaluators, and why both
- ConversationCompleted is deterministic: did the assistant answer every scripted turn? It runs first because any other claim about a conversation that stopped at turn one of three is a claim about a different conversation. A truncated conversation fails it; it is a result, not a missing case.
- ConversationJudge is a model judge over the whole transcript, every user and assistant turn, because the failures a conversational product is blamed for are relational: an answer that contradicts one from the turn before is wrong only beside it. A truncated conversation is judged on what was recorded, and the transcript tells the judge where it stopped.
evaluators = [
ConversationCompleted(),
ConversationJudge(criterion="plan_coherent", provider=..., model=..., rubric_text=RUBRIC),
]Judging offline
A judge needs a model. To run with no network, judge_offline.py passes a scripted provider, the same helper Oloproof's own tests use (FakeProvider from the engine internals, not public API). It answers every judge prompt by one fixed rule: pass when every assistant turn names the plan the reference names. That makes the judgments deterministic and the tutorial reproducible. It measures nothing about how a real judge model behaves.
The policy
release.yaml:
version: 1
confidence_level: 0.95
block_on: [FAIL, INSUFFICIENT_EVIDENCE, MANUAL_REVIEW]
rules:
- {id: completion-floor, metric: conversation_completed, min: 0.75}
- {id: coherence-floor, metric: plan_coherent, min: 0.80}Run it
python evaluate.pybaseline run run_...
conversation_completed: 0.900 [0.763, 0.972] over 40 conversations
plan_coherent: 0.400 [0.249, 0.567] over 40 conversations
gate BLOCK (exit 3)
completion-floor: PASS (lower_bound_meets_minimum)
coherence-floor: INSUFFICIENT_EVIDENCE (evaluator_not_validated)
first failing conversation:
conversation_completed: passed=True {'turns_recorded': 2, 'turns_declared': 2, 'truncated': False}
plan_coherent: passed=False {'provider_model': 'offline-rule'}How to read it:
- Completion is 36 of 40, the four hand-offs. Its lower bound clears 0.75, so that rule is PASS.
- Coherence is 40%, and its rule reads INSUFFICIENT_EVIDENCE with evaluator_not_validated, not FAIL. A judge's rules do not decide until the judge has been measured against human labels (require_validated_evaluators is on by default; Judges explains it). The estimate is still shown, and it is still evidence: it just cannot release or block on its own.
- gate BLOCK (exit 3): the policy blocks on INSUFFICIENT_EVIDENCE. Exit 3 is that state, not a failure.
Validating the offline stand-in would be meaningless, since it is a rule written for this example. With a real judge, label a sample of the run with oloproof review RUN_ID --criterion plan_coherent --by YOU --sample 20 and then run oloproof evaluators validate EVALUATOR_ID --by YOU.
Inspect a failing conversation
The SDK writes to the same store the CLI reads, .oloproof/ in the directory you ran from:
oloproof inspect RUN_ID --case conv_00output: {
"answer": "The starter plan does not include SSO.",
"turns_answered": 2
}
judgments:
conversation_completed: passed
plan_coherent: failed
judge text, not verified:
every answer is about enterprise: FalseThe user asked about the enterprise plan and the follow-up was answered about the starter plan. oloproof inspect RUN_ID --failures lists every failing conversation; all 24 follow-ups without a plan name fail the same way. The next action is in the application: keep the plan in session state.
A candidate change, and the comparison
candidate in systems.py sets remembers_plan=True. evaluate.py runs both systems over the same scripts and compares them case by case under comparison.yaml:
rules:
- {id: coherence-better, kind: superiority, metric: plan_coherent}
- {id: completion-no-worse, kind: non_inferiority, metric: conversation_completed, margin: 0.05}comparison, candidate minus baseline
conversation_completed: +0.000 [-0.127, +0.127] over 40 pairs
plan_coherent: +0.600 [+0.337, +0.817] over 40 pairs
gate BLOCK (exit 3)
coherence-better: INSUFFICIENT_EVIDENCE (evaluator_not_validated)
completion-no-worse: INSUFFICIENT_EVIDENCE (interval_overlaps_margin)The coherence difference is large and its interval excludes zero, but the judge is unvalidated, so its rule still does not decide. Completion is unchanged, and 40 pairs cannot show it is within five points: the interval reaches 12.7 points either way. Both point to the same next step: validate the judge and add conversations.
The alternative: turns as cases, analysed as clusters
A conversation-level verdict says how often a conversation went well, not which turn went wrong. Since Oloproof has no turn-level metric within a recorded conversation, the other route is to make each turn its own case and tie a conversation's turns together with group_id:
{"expected": {"plan": "enterprise"}, "group_id": "conv_00", "id": "conv_00_t2", "input": {"history": ["What does the enterprise plan cost?"], "question": "Does that include SSO?"}}The system replays the scripted history into a fresh session, then answers the turn:
@system(name="plan-assistant-turns", version="baseline")
def turn_baseline(case):
session = PlanAssistant(remembers_plan=False)
for earlier in case["history"]:
session.ask(str(earlier))
return session.ask(str(case["question"]))Turns of one conversation are not independent, so once any case has a group_id the suite is analysed by cluster with an approximate method a policy must accept (turns_release.yaml sets allow_approximate_methods: true; Clustered cases explains it):
turns as cases: turn_plan 0.667 [0.588, 0.749] over 72 turns
gate BLOCK (exit 1)
turn-plan-floor: FAIL (upper_bound_below_minimum)The trade-off:
| One case per conversation | One case per turn, clustered | |
|---|---|---|
| Unit of the rate | conversations that went well | turns answered correctly |
| Effective sample size | the number of conversations | still the number of conversations, not turns |
| Which turn failed | read the transcript | each turn has its own verdict |
| History each turn sees | the assistant's own earlier answers | the script's earlier user turns, replayed |
| Catches drift caused by its own earlier answers | yes | no, each turn starts from a scripted history |
| Evaluators | ConversationCompleted, ConversationJudge (SDK only) | any evaluator, in YAML or the SDK |
The turns alternative excludes the four hand-off conversations, so its 72 turns come from 36 conversations. Here it can decide where the judge could not, because ExactMatch is deterministic and needs no validation.
Optional: a live judge model
This step needs a model served on your machine. It is not run by the offline tutorial or its test. With Ollama running and llama3.1 pulled:
python evaluate.py --liveThe judge then calls http://localhost:11434/v1 with provider="openai_compatible". A loopback server needs no key and sends nothing off the machine. A cloud provider needs its key in the environment, sends each transcript to that provider and costs money per judgment. A real model's verdicts differ from the stand-in's, so the numbers above will change, and its rules still read evaluator_not_validated until you validate it.
Troubleshooting
| Symptom | Cause and fix |
|---|---|
| this evaluator needs exactly one conversation/v1 artifact; the case recorded 0 | The adapter did not call current_case().artifact(CONVERSATION, ...), or raised before it. Record even when the conversation stops early. |
| malformed conversation/v1 artifact: ... 0 of 2 turns recorded and truncated is false | The run stops with a SystemContractError. A recording with fewer turns than declared_turns must set truncated: true. |
| conversation turn indexes must be contiguous starting at 1 | Number turns 1, 2, 3 by the scripted turn they answer. |
| evaluator 'conversation_completed' needs conversation/v1 artifacts, but system ... does not declare that it records them | Add records=(CONVERSATION,) to the @system decorator. |
| Input tag 'conversation_completed' found using 'type' does not match any of the expected tags from oloproof run | Conversation evaluators are SDK only. Use a script as here. |
| Answers leak between conversations | A session is shared between cases. Create one per case. |
| Coherence rules never decide | The judge is unvalidated. Validate it, or set require_validated_evaluators: false knowingly. |
Limitations
- No user simulator: every user turn comes from the dataset script, so the conversation cannot branch on what the assistant said.
- No turn-level metric within a recorded conversation; use turns as clustered cases instead, with the trade-off above.
- No conversation replay: a recorded conversation cannot be re-run against another system. Comparing two systems means each one replays the same script.
- ConversationCompleted and ConversationJudge are Python SDK only.
- Text only: a judge is shown JSON text, never images or audio.
- The offline judge is a scripted rule. Its verdicts show the mechanics, not a real judge's accuracy.
Where to go next
- Judges covers providers, validation and recalibration.
- Clustered cases covers group_id and the approximate-method opt-in.
- The Python API covers evaluate and evaluate_comparison.
- Agents covers the same boundary for an agent's tool loop.