VoiceHaul evaluation and diagnosis for empathic voice agents
Vahit Feryad, PhD · source · full report

A turn-level rating tells you whether a reply sounded right. It cannot tell you whether the agent still honours what the caller asked for twenty turns ago — or how much of your rating panel an LLM judge could actually replace.

Pick a tab, set the inputs on the left, press the button. Results appear on the right. Every tab already shows a worked example, nothing calls an external service, and the harness is pure Python.

Your inputs

which quality is being rated
which callers

The judge and the model being graded here are real; the callers are simulated so the true quality of each turn is known and the estimator can be checked.

What you get

1 judge rating is worth
0.01
human ratings — poor
95% CI 0.00 to 0.03
judge reliability
0.01
rho on n=160 turns — poor
one human rater: 0.71
saving per year
$0
not quotable — the interval does not clear the floor

The panel cannot be replaced on this dimension. The judge is worth 0.01 human ratings here, below the 0.25 floor. There is nothing to save on this dimension: it stays on the panel.

panel only what the point estimate would imply
conversations per release 148 148
human ratings 444 444
judge ratings 0 0
cost per release $1,332 $1,332

Sized for a suite that can detect a 20% regression. The ratio comes from the Spearman-Brown formula inverted, and it expires whenever the judge model, the domain or the rubric changes.

Every segment at once. Anything left of the dashed line cannot substitute.

voicehaul-judge-substitution .json 45.4 KB ⇣
one turn caller + reply human panel k raters, noisy LLM judge cheap, unproven latent truth simulation only correlate, then disattenuate divide out panel noise estimated rho what a customer gets is the estimate right? this arrow does not exist on real data: both measurements carry error and neither one is the reference
The estimator a customer can run uses only the two solid paths. The dashed path is the validation, and it is the reason this is built in a simulator before it is pointed at a contract.
dimension n judge reliability reading 1 judge = N human reading
perceived empathy 160 0.18 poor 0.29 weak
did it actually help 160 0.01 poor 0.01 poor

The estimator, checked against ground truth — the one validation real data cannot run:

dimension estimated rho true rho error
perceived empathy 0.178 0.189 -0.011
did it actually help 0.014 0.016 -0.002

Judge: meta-llama/llama-3.1-8b-instruct. A single agreement figure would have reported 0.18 and hidden a twenty-fold spread across segments.

Each dot is one turn. Horizontal: the quality the turn actually had; vertical: what each rater said. Both lines slope the same way — the judge is not backwards. What separates them is the spread, and spread is what reliability measures.

A metric without a band is a number nobody can act on. But half of these cannot honestly carry a fixed threshold, and pretending otherwise is how a dashboard starts lying.

  • Absolute — means the same thing on any suite.
  • Anchored — a distance from an ideal defined for one action space, read against the range the suite can actually reach.
  • Conventional — thresholds from psychometrics, cited rather than invented.
feedback uptake @10
absolute
Share of the caller's explicit requests still honoured ten turns later. 1.00 means every request survives the call.
good ≥ 0.9  ·  acceptable ≥ 0.7  ·  weak ≥ 0.45
Absolute: a proportion of requests, identical in meaning on any suite.
conversation failure rate
absolute
Share of calls that ended with the caller worse off than they started. Anything above one in five is a support problem, not a measurement artefact.
good ≤ 0.05  ·  acceptable ≤ 0.2  ·  weak ≤ 0.45
Absolute: the threshold is the outcome definition in env/episode.py.
left-over distress
absolute
Negative affect the caller is still carrying over the last five turns, on a 0-1 scale. A call is counted as failed above 0.55.
good ≤ 0.25  ·  acceptable ≤ 0.45  ·  weak ≤ 0.6
Absolute: same scale as the failure threshold, so the two agree.
calibration
anchored
How close each turn came to the best available response for the caller's state. Read against the range this suite can actually reach, never as an absolute score.
good ≥ 0.85  ·  acceptable ≥ 0.6  ·  weak ≥ 0.35
Anchored: the ideal policy is defined over one action space, so the raw number is not comparable across systems. Floor and ceiling are measured from the suite.
panel: perceived empathy
anchored
What a rater scoring one held-out turn in isolation would reward. Higher is not automatically better here - this is the number that can move the wrong way.
good ≥ 0.85  ·  acceptable ≥ 0.6  ·  weak ≥ 0.35
Anchored, and deliberately not a target: it is the contrast.
judge reliability
conventional
How much of the real variation in quality the automated judge captures. 0.80 is the usual bar for trusting a measure on individual cases; below 0.40 it carries too little signal to act on.
good ≥ 0.8  ·  acceptable ≥ 0.6  ·  weak ≥ 0.4
Conventional: the 0.70-0.80 range is the long-standing threshold for reliability in psychometrics (Nunnally); 0.40 is where a measure stops separating cases at all.
substitution ratio
conventional
How many human ratings one judge rating is worth. At 1.00 the judge replaces a rater one for one; below 0.25 it is not worth the pipeline.
good ≥ 1.0  ·  acceptable ≥ 0.5  ·  weak ≥ 0.25
Derived: the Spearman-Brown formula inverted, so the cut points are the reliability cut points above expressed as raters.
caller state 6-dim affect renders utterance what is said model under test reply delivery vector rate, warmth, length, apology, acknowledgement perceived empathy calibration hidden what a rater scores what moves the caller drives the next state feeds nothing
One turn. The same delivery vector is scored twice, and only one of the two changes what happens next. The observable score is the one with no arrow leaving it — which is why a turn-level suite and a conversation can disagree.

Every reply is scored twice. Perceived empathy is what a rater scoring one held-out turn in isolation would reward. Calibration is how close the reply came to the best available response for the caller's actual state, and it is hidden from the rater. Where those two diverge is the finding.

Five reference policies each embody exactly one known failure mode; five caller personas each need a different kind of handling. Soothing requires delivering below the caller's energy, so a policy that mirrors warmth back scores well on every single turn and loses the conversation.

Simulated: the callers, the five reference policies, the affect dynamics.

Real: the metrics, the estimators, the statistics, the localization algorithm — and, in the first tab, the model being graded and the judge grading it.

The reason for a synthetic environment is not convenience. You cannot validate a measurement instrument without ground truth you control. If you only run an eval against real models, a metric that reports the wrong thing and a model that behaves badly are indistinguishable.


Pointing it at a real model exposed seven faults, every one of them in the instrument rather than the model:

  1. An unmeasured speech rate was scored as if measured, injecting a fabricated error and making soothing structurally impossible.
  2. The acknowledgement pattern matched I understand but not I can understand, which is what instruction-tuned models write.
  3. The fixed-context panel gave a speaking policy no utterance, so it measured a real model's reply to an empty turn.
  4. Both panel dimensions replayed each block separately, doubling the cost of every hosted call.
  5. The break-even calibration was anchored on simulated policies and sat at the 88th percentile of what a real model reaches, pinning every conversation to the failure floor. It is now a measured parameter recorded in the suite id, and the gate detects that saturation and refuses to report a meaningless delta.
  6. A scatter caption claimed the judge was flat. It is not; the slopes agree and the spread does not.
  7. An entire tab was written and never rendered.

Every one has a regression test. An evaluation harness that is wrong is worse than no harness, because it is believed.

Four more instruments run in the repository and are written up in the full report:

  • Watch a conversation — one call turn by turn, with the turn that broke it marked and named.
  • Which turn broke it — failure-onset localization, 97% top-1 against injected faults with no false positives on healthy calls.
  • Rating budget — the smallest regression a suite could detect, and which conversations to spend the budget on.
  • Calibrate — anchoring the break-even point on measured behaviour instead of an assumption.

Full report · Source · Runopsy · LongHaul-Bench · Apache-2.0

built by
Vahit Feryad, PhD
Applied AI research engineer · evaluation, benchmarking and agent reliability
PhD in electrical and electronics engineering · 10+ years industrial R&D · 237 Google Scholar citations · Istanbul
the work this is built on
LongHaul-Bench — long-horizon agent reliability over 1,000+ sequential episodes, five-world programme, memory ablations
Runopsy — causal failure-onset diagnosis with counterfactual replay (Apache-2.0, on PyPI)
source  ·  scholar  ·  linkedin  ·  static report