voicehaul · long-horizon evaluation for empathic voice agents
A turn-level rating tells you whether a response sounded right. It cannot tell you whether a model still honours what the user asked for twenty turns ago, whether its calibration decays as a session runs long, whether it is regulating the user's affect or mirroring it back, or which turn broke a call that ended badly.
This measures those four things, then measures how much rating budget you need before any of them is detectable. Runs in about fifteen seconds with no API key and no network.
Provenance
The measurement design here is not new work. It is the method from LongHaul-Bench — a long-horizon reliability benchmark run over 1,000+ sequential episodes with a five-world experimental programme and memory-strategy ablations — and Runopsy, a causal failure-onset method using deterministic diagnostics plus counterfactual replay. Both are carried across here from text agents to voice. Building this instance took a day because the underlying method took two years. What follows is phase zero: the synthetic validation step that has to pass before any of it is pointed at real audio. The section at the bottom is the part that matters.
The finding
On the left, a fixed-context turn panel: every agent answers the same held-out user states and each turn is rated on its own, the way a prompt-set leaderboard works. On the right, the same agents actually holding forty-turn conversations.
The turn panel's best model is the one that fails every conversation. Rank correlation between the two orderings: .
Not because turn-level rating is wrong — it is the right instrument for the question it asks. It is because a turn panel holds the context fixed, and in a real conversation the model creates the context it is later scored on. A model that keeps users agitated is subsequently asked easier-looking questions.
Explore
Pick a policy and a caller. The trace shows what the user is carrying, what a turn-level rater would score, and what the turn was actually worth. The log below marks every explicit request and every new grievance.
Put mirror against the hostile caller. Perceived empathy stays respectable the whole way down — it sounds attuned on every single turn — while distress never comes down. It matches the caller's energy instead of sitting just below it, so the caller has nowhere to come down to. No turn is bad. The conversation is.
Then put drifter against anyone and watch what happens a dozen turns after a request: it complies for a while, then quietly stops, and speech rate creeps back up. The turn-level score barely moves.
Metric one
Users voice explicit requests — slow down, stop apologising, be concise. Uptake is the share still honoured some number of turns later. Two policies with the same turn-level score can sit at opposite ends of this curve.
Metric two
Two readings of the same behaviour. Horizontally: how closely the agent's vocal energy tracks the user's — how attuned it sounds. Vertically: how much negative affect it actually removes per turn. A single "empathy" score collapses these into one number and loses the distinction that matters.
Metric three
A share of user requests reach the model as the opposite instruction — an ASR error, or an adversarial user. The user still expects the original.
The agents that listen best degrade most. Perfect compliance with a corrupted channel is itself a failure mode, and it is invisible to any suite whose feedback channel is assumed clean.
Diagnosis
A regression is injected at a turn the diagnostic is never told about: the model loses its conditioning and reverts to a generic upbeat persona. Four cheap deterministic signals propose candidates, a walk-back finds where the anomalous stretch begins, and counterfactual replay gates — if repairing the agent from the proposed turn would not have changed the outcome, it returns no answer rather than a wrong one.
Lower the severity to blend the fault into the model's own policy. That is the realistic case, and the honest one to report.
Budget
Human ratings are both the ground truth and the budget line. Between-conversation spread is measured from the suite; per-rater noise is yours to set. This is the calculation that turns "we track regressions" into a number of conversations and a cost.
Read this part carefully
Simulated: the users, the agents, and the affect dynamics. The five agents are deterministic policy simulators, not language models — each embodies exactly one known failure mode. The user model and both scoring functions are written by hand.
Real: the metrics, the estimators, the statistics, and the localization algorithm. Those are the deliverable.
The reason for a synthetic environment is not convenience. You cannot validate a measurement instrument without ground truth you control. If you only ever run an eval against real models, a metric that reports the wrong thing and a model that behaves badly are indistinguishable. Here the fault turn is known, the failure mode is known, and the ideal policy is known — so "93% accurate at severity 0.5 with no false positives" is a checkable statement about the method, not a leaderboard entry.
The rank-correlation result is the one to read sceptically. I encoded the hypothesis that turn-level raters reward attunement and warmth while outcomes reward down-regulation, and the environment then confirms it — which is close to circular on its own. Its value is that the hypothesis is now falsifiable against real rater data: fit the perceived-empathy function to real human ratings, keep the outcome measure, and the same code reports whether the gap survives. If it doesn't, that is a genuinely useful negative result about a class of eval suites.
The design property that makes this practical: the harness needs no privileged access to the model under test. Both sides of the conversation are scored from audio — the user's affect trajectory from expression measurement on the user channel, and the agent's delivery parameters from expression measurement plus transcript statistics on the agent channel. The same metrics apply to a model you own, one you licence, and a competitor's public endpoint.
Two things I would fix before trusting it on real audio, in order. The compact affect basis is currently a hand-written alias map over a 48-category readout; that projection should be fitted against human ratings of the same clips, because everything downstream inherits its error. And apology and acknowledgement detection uses English keyword lists; for a multilingual suite that has to become a small classifier. The prosodic features generalise, the lexical ones do not.
The commercial question
A human panel is the trusted measurement and the dominant cost line. An LLM judge is cheap and of unknown trustworthiness. Every voice-evaluation contract runs into the same question, and the honest answer today is “sometimes, and we cannot tell you when”.
That answer is not one number. It is a number per dimension and per caller segment: a judge can track how empathic a turn sounds and be near-blind to whether it actually helped.
| dimension | n | judge ρ estimated |
1 judge rating = N human | verdict |
|---|---|---|---|---|
| perceived empathy | 160 | 0.18 | 0.29 | supplement |
| did it actually help | 160 | 0.01 | 0.01 | human only |
A single agreement figure would have reported 0.18 and hidden a twenty-fold spread across segments. On whether a turn actually helped, the judge scores 0.00 with hostile callers — blind precisely on the calls that generate escalations.
perceived empathy, by caller
| caller | judge ρ | = N human |
|---|---|---|
| confused elderly | 0.07 | 0.04 |
| distressed billing | 0.16 | 0.11 |
| hostile escalation | 0.14 | 0.21 |
| cautious optimist | 0.20 | 0.28 |
| grieving claim | 0.37 | 1.56 |
did it actually help, by caller
| caller | judge ρ | = N human |
|---|---|---|
| hostile escalation | 0.00 | 0.00 |
| distressed billing | 0.01 | 0.00 |
| grieving claim | 0.01 | 0.00 |
| confused elderly | 0.02 | 0.01 |
| cautious optimist | 0.05 | 0.01 |
Standard psychometrics, not a new idea. The reliability of one human rating, the Spearman–Brown formula for a panel of k, and the correlation between judge and consensus corrected for that consensus’s own unreliability. Inverting Spearman–Brown gives the substitution ratio:
k* = ρg(1 − ρh) / [ρh(1 − ρg)]
A customer estimates the judge’s reliability by correlating it against a human consensus and dividing out that consensus’s unreliability. Whether that estimate is right is unknowable in the field: both measurements carry error and neither is the reference. Here the latent quality is known by construction.
| dimension | estimated ρ what a customer computes |
true ρ known here | error |
|---|---|---|---|
| perceived empathy | 0.178 | 0.189 | -0.011 |
| did it actually help | 0.014 | 0.016 | -0.002 |
Note what does not go away: estimating the judge’s reliability requires human ratings, and the estimate expires whenever the judge model, the domain or the rubric changes. This is a recurring measurement, not a setting.
The programme
Everything above is the part that can be done in a synthetic environment, and it is the smallest part. It establishes that the instrument measures what it claims to measure on a case where the answer is known by construction. That is a precondition, not a result.
The four questions below are the actual research. Each is a quarter or more of work, each is gated on data that does not exist yet, and each produces both a publishable finding and a product surface. They are listed with their cost rather than their promise, because a benchmark that is oversold once is never trusted again.
Why it matters. The headline result rests on a hypothesis encoded by hand: that turn-level raters reward attunement and warmth while conversation outcomes reward down-regulation. In a synthetic world that is close to circular. Against real ratings it becomes falsifiable, and it is the one claim everything else depends on.
What answering it takes. Several hundred real conversations, each rated twice — turn by turn in isolation, and once at the conversation level — with multiple raters per item, because the power table above is what says how many. This is a rating-panel study, not a code change.
What it produces. Either a validated instrument, or a clean negative result about a whole class of evaluation suites. Both are worth publishing; one of them is worth building a product on.
One to two quarters, almost all of it data collection and inter-rater reliability work
Why it matters. Down-regulation is a cultural norm as much as an acoustic one. Speaking below someone's energy may not soothe identically in every language, and the lexical features here are English-only by construction. If the instrument is not invariant, a cross-language leaderboard compares different quantities and calls it one score. For a fifty-language product that is not a footnote.
What answering it takes. Parallel evaluation corpora across several typologically distinct languages, native-speaker rater panels, and formal measurement-invariance testing — configural, metric, scalar.
What it produces. A methods paper of the kind cited by everyone who later builds a multilingual voice benchmark, and a defensible answer to the first question a serious enterprise buyer asks.
Two to three quarters, running partly in parallel with 01
Why it matters. A 48-category expression readout is collapsed onto a six-dimension basis by an alias map written by hand. Every metric downstream inherits that projection's error and nobody currently knows how large it is. This is the least glamorous item here and probably the highest-leverage one.
What answering it takes. Paired data — expression measurement and human ratings on the same clips — then fitting the projection against the ratings rather than asserting it, with held-out validation.
What it produces. A measurable accuracy improvement across every metric at once, and a reusable component rather than a bespoke one.
A quarter once the data from 01 exists; it cannot start before
Why it matters. The real test of an evaluation is not whether it correlates with human judgement. It is whether using it as a training signal moves a model on held-out human judgement. This is where evaluation stops being measurement and becomes post-training, and it is the only one of these four that changes what a model is rather than what we know about it.
What answering it takes. Everything above, plus a model that can actually be fine-tuned and a held-out human evaluation that was never used for training.
What it produces. The finding that would justify the whole programme, or the finding that these metrics are diagnostic but not optimisable — which is itself important and rarely reported.
Four quarters or more, and it should not be started early
None of these four compresses by working harder: three are gated on rating data that has to be collected, and the fourth is gated on the other three. That is the shape of the work — not a tool to be delivered, but an instrument that earns authority by being maintained, versioned, contamination-checked and re-validated as the models it measures keep moving. A benchmark nobody maintains stops being cited in about a year.
The two projects this one is built on took about two years between them. LongHaul-Bench ran a five-world experimental programme over a thousand-plus sequential episodes per run under hard edge constraints, comparing frozen, append-only, reflect, gated and oracle memory strategies with full ablations and statistical analysis. Runopsy is the causal failure-onset method that the localization section above is a port of.
That is why this instance came together quickly, and it is also why it would be wrong to call it quick work. The fast part is the part that was already solved. What is left — validating a measurement instrument against human judgement, establishing that it holds across languages, and finding out whether it can be optimised against — is ordinary research on an ordinary research timescale.