A turn-level rating tells you whether a reply sounded right. It cannot tell you whether the agent still honours what the caller asked for twenty turns ago — or how much of your rating panel an LLM judge could actually replace.
Pick a tab, set the inputs on the left, press the button. Results appear on the right. Every tab already shows a worked example, nothing calls an external service, and the harness is pure Python.
Your inputs
The judge and the model being graded here are real; the callers are simulated so the true quality of each turn is known and the estimator can be checked.
What you get
95% CI 0.00 to 0.03
one human rater: 0.71
The panel cannot be replaced on this dimension. The judge is worth 0.01 human ratings here, below the 0.25 floor. There is nothing to save on this dimension: it stays on the panel.
| panel only | what the point estimate would imply | |
|---|---|---|
| conversations per release | 148 | 148 |
| human ratings | 444 | 444 |
| judge ratings | 0 | 0 |
| cost per release | $1,332 | $1,332 |
Sized for a suite that can detect a 20% regression. The ratio comes from the Spearman-Brown formula inverted, and it expires whenever the judge model, the domain or the rubric changes.
Every segment at once. Anything left of the dashed line cannot substitute.
| voicehaul-judge-substitution .json | 45.4 KB ⇣ |
| dimension | n | judge reliability | reading | 1 judge = N human | reading |
|---|---|---|---|---|---|
| perceived empathy | 160 | 0.18 | poor | 0.29 | weak |
| did it actually help | 160 | 0.01 | poor | 0.01 | poor |
The estimator, checked against ground truth — the one validation real data cannot run:
| dimension | estimated rho | true rho | error |
|---|---|---|---|
| perceived empathy | 0.178 | 0.189 | -0.011 |
| did it actually help | 0.014 | 0.016 | -0.002 |
Judge: meta-llama/llama-3.1-8b-instruct. A single agreement figure would have reported 0.18 and hidden a twenty-fold spread across segments.
Each dot is one turn. Horizontal: the quality the turn actually had; vertical: what each rater said. Both lines slope the same way — the judge is not backwards. What separates them is the spread, and spread is what reliability measures.
Your input
Paste a transcript. Prefix lines with Caller: and Agent:. The box starts with a worked example — replace it with one of yours.
Only what a transcript can honestly answer is reported: what the agent's delivery was, whether the caller asked for a change, and whether the agent did it and kept doing it. Nothing is inferred that you could not check by hand.
What you get
Speaker labels found and used.
| request | asked | honoured | first 3 turns |
|---|---|---|---|
| stop apologizing | 1 | 1/3 | 1/3 |
| be concise | 1 | 1/3 | 1/3 |
Turns worth looking at
- turn 3: ignored: stop apologizing, be concise
- turn 7: ignored: stop apologizing, be concise
What is deliberately not reported: whether the caller ended up better off. A transcript carries no ground truth about how they felt, and guessing it would be the kind of number that gets believed. On audio, expression measurement supplies it.
| voicehaul-your-call .csv | 1.1 KB ⇣ |
Your inputs
Try calibrated against mirror: the panel rates the candidate higher and the gate blocks it anyway.
What are these five policies?
What you get
- calibration regressed by -0.123 (Holm p=0.0000)
- the fixed-context turn panel rated the candidate HIGHER on perceived empathy - a leaderboard would have passed this release
| dimension | baseline | candidate | delta | holm p |
|---|---|---|---|---|
| what a fixed-prompt leaderboard reports | ||||
| panel: perceived empathy diagnosticacceptable floor 0.00ceiling 1.00 | 0.577 | 0.635 | +0.058 | 0.0000 |
| panel: calibration diagnostic | 0.941 | 0.795 | -0.145 | 0.0000 |
| what the conversations report | ||||
| calibrationacceptable floor 0.00ceiling 1.00 | 0.965 | 0.842 | -0.123 | 0.0000 |
| feedback uptake @10good | 1.000 | 1.000 | +0.000 | 1.0000 |
| left-over distressgood | 0.104 | 0.194 | +0.090 | 0.2751 |
| conversation failure rateacceptable | 0.000 | 0.133 | +0.133 | 0.1040 |
| calibration drift diagnostic | 0.564 | 5.805 | +5.242 | 0.0000 |
Where it lands — calibration by caller
Suite 59b315bde1ce | 30 conversations × 40 turns × 5 callers. Welch two-sample tests with one conversation as the unit, Holm-corrected across 7 dimensions.
Diagnosis. 4 of 4 failed candidate conversations were localized to an onset turn; median turn 5. Most common cause: affect escalation (3), delivery discontinuity (1).
What this suite could not see. At n=30 with 3 raters, the smallest detectable regression is 0.44 Likert points; resolving a 0.20-point change needs n = 148.
The gate exits non-zero on BLOCK, so it gates a release without any glue code. The JSON above is the same report a job would parse.
# .github/workflows/voice-eval.yml
- name: Long-horizon regression gate
run: |
pip install -e .
voicehaul gate $BASELINE $CANDIDATE \
--config configs/support-en-40turn.yaml \
--format json --out artifacts/
# exit 1 = a gating dimension regressed; the job fails and the build stops
- uses: actions/upload-artifact@v4
with:
name: voice-eval
path: artifacts/
Every run is stamped with a suite id derived from the config, so two reports carrying the same id were measured the same way and two carrying different ids were not.
A metric without a band is a number nobody can act on. But half of these cannot honestly carry a fixed threshold, and pretending otherwise is how a dashboard starts lying.
- Absolute — means the same thing on any suite.
- Anchored — a distance from an ideal defined for one action space, read against the range the suite can actually reach.
- Conventional — thresholds from psychometrics, cited rather than invented.
| feedback uptake @10 absolute | Share of the caller's explicit requests still honoured ten turns later. 1.00 means every request survives the call. good ≥ 0.9 · acceptable ≥ 0.7 · weak ≥ 0.45 Absolute: a proportion of requests, identical in meaning on any suite. |
| conversation failure rate absolute | Share of calls that ended with the caller worse off than they started. Anything above one in five is a support problem, not a measurement artefact. good ≤ 0.05 · acceptable ≤ 0.2 · weak ≤ 0.45 Absolute: the threshold is the outcome definition in env/episode.py. |
| left-over distress absolute | Negative affect the caller is still carrying over the last five turns, on a 0-1 scale. A call is counted as failed above 0.55. good ≤ 0.25 · acceptable ≤ 0.45 · weak ≤ 0.6 Absolute: same scale as the failure threshold, so the two agree. |
| calibration anchored | How close each turn came to the best available response for the caller's state. Read against the range this suite can actually reach, never as an absolute score. good ≥ 0.85 · acceptable ≥ 0.6 · weak ≥ 0.35 Anchored: the ideal policy is defined over one action space, so the raw number is not comparable across systems. Floor and ceiling are measured from the suite. |
| panel: perceived empathy anchored | What a rater scoring one held-out turn in isolation would reward. Higher is not automatically better here - this is the number that can move the wrong way. good ≥ 0.85 · acceptable ≥ 0.6 · weak ≥ 0.35 Anchored, and deliberately not a target: it is the contrast. |
| judge reliability conventional | How much of the real variation in quality the automated judge captures. 0.80 is the usual bar for trusting a measure on individual cases; below 0.40 it carries too little signal to act on. good ≥ 0.8 · acceptable ≥ 0.6 · weak ≥ 0.4 Conventional: the 0.70-0.80 range is the long-standing threshold for reliability in psychometrics (Nunnally); 0.40 is where a measure stops separating cases at all. |
| substitution ratio conventional | How many human ratings one judge rating is worth. At 1.00 the judge replaces a rater one for one; below 0.25 it is not worth the pipeline. good ≥ 1.0 · acceptable ≥ 0.5 · weak ≥ 0.25 Derived: the Spearman-Brown formula inverted, so the cut points are the reliability cut points above expressed as raters. |
Every reply is scored twice. Perceived empathy is what a rater scoring one held-out turn in isolation would reward. Calibration is how close the reply came to the best available response for the caller's actual state, and it is hidden from the rater. Where those two diverge is the finding.
Five reference policies each embody exactly one known failure mode; five caller personas each need a different kind of handling. Soothing requires delivering below the caller's energy, so a policy that mirrors warmth back scores well on every single turn and loses the conversation.
Simulated: the callers, the five reference policies, the affect dynamics.
Real: the metrics, the estimators, the statistics, the localization algorithm — and, in the first tab, the model being graded and the judge grading it.
The reason for a synthetic environment is not convenience. You cannot validate a measurement instrument without ground truth you control. If you only run an eval against real models, a metric that reports the wrong thing and a model that behaves badly are indistinguishable.
Pointing it at a real model exposed seven faults, every one of them in the instrument rather than the model:
- An unmeasured speech rate was scored as if measured, injecting a fabricated error and making soothing structurally impossible.
- The acknowledgement pattern matched I understand but not I can understand, which is what instruction-tuned models write.
- The fixed-context panel gave a speaking policy no utterance, so it measured a real model's reply to an empty turn.
- Both panel dimensions replayed each block separately, doubling the cost of every hosted call.
- The break-even calibration was anchored on simulated policies and sat at the 88th percentile of what a real model reaches, pinning every conversation to the failure floor. It is now a measured parameter recorded in the suite id, and the gate detects that saturation and refuses to report a meaningless delta.
- A scatter caption claimed the judge was flat. It is not; the slopes agree and the spread does not.
- An entire tab was written and never rendered.
Every one has a regression test. An evaluation harness that is wrong is worse than no harness, because it is believed.
Four more instruments run in the repository and are written up in the full report:
- Watch a conversation — one call turn by turn, with the turn that broke it marked and named.
- Which turn broke it — failure-onset localization, 97% top-1 against injected faults with no false positives on healthy calls.
- Rating budget — the smallest regression a suite could detect, and which conversations to spend the budget on.
- Calibrate — anchoring the break-even point on measured behaviour instead of an assumption.
Full report · Source · Runopsy · LongHaul-Bench · Apache-2.0
Runopsy — causal failure-onset diagnosis with counterfactual replay (Apache-2.0, on PyPI)