Three Humans Heard a Safe Handoff. Two Models Scored It a Failure.
A Sprinklr AI paper scored 242 voice-agent conversations with three human annotators and two OpenAI judges, and found the judges rank retail service metrics reliably while recording a fraction of the safety behaviour humans see. The gap lands on the one call a retailer most needs graded correctly.
Neritus Vale
A retailer running voice agents on returns and order changes cannot listen to the calls, so a second model listens and assigns the grade. Which of those grades hold up is the question nobody has priced. A paper revised on arXiv yesterday by researchers at Sprinklr AI scored 242 voice-agent conversations twice: once by three trained human annotators, once by GPT-4.1 and GPT-5 on the same rubric. The two evaluators sort the metrics the same way: what scores high for a human scores high for the judge. They disagree on the level, and on the safety metrics that disagreement reaches a factor of six. The paper’s own recommendation: LLM evaluators are “best used as scalable first-pass tools.”
The sample is narrower than the headline count suggests. Retail supplied 120 of those conversations, and 98 of them were returns, exchanges, order modifications or cancellations. That is the transactional end of service, where the right answer is checkable against an order record. Clienteling is absent: no styling call, no sizing argument, no customer whose gift did not arrive. The finding still lands, because Sprinklr sells what it is testing: quality-management software that “automatically scores 100% of customer interactions across voice, chat, email, social, and messaging channels.” A vendor publishing the metric-level limits of automated scoring is arguing against its own product page.
Retail is where the judge looks good, and the reason is the domain rather than the model. Across the ten metrics, human and judge scores in retail correlate at 0.912 for GPT-4.1. The authors read that as a judge capturing the right shape and needing only a rescale, a per-metric linear correction fitted once. In telecom the same correlation falls to 0.295 and stops being statistically significant. Same judges, same rubric, same three prompting configurations; what changes is the conversation.
The exception is safety, and it is not a rounding difference. Two metrics carry it: Safety Recall, how consistently the agent asks for confirmation when it should, and Irreversible Action Safety, the share of cancellations and charges executed only after the customer agreed. In telecom, all three annotators put Irreversible Action Safety between 0.794 and 1.000, which describes an agent that nearly always confirms before acting. The judges, reading the same transcripts under the same rubric, land near a fifth of that. The paper treats a shortfall on that metric as a critical safety failure, so the two evaluators are not describing the same call.
The authors trace the gap to one ambiguity, and it is the ambiguity retail service turns on. When an agent hands a high-risk action to a person, the judge swings between two readings of identical behaviour: crediting the agent for escalating, or penalising it for not finishing the job. Rubric text for a necessary escalation, they write, “is hard to state without ambiguity.” The inconsistency that follows “suppresses the recall of real safety events.” A human annotator applies situational judgment to the transcript; a judge scoring in one pass has nothing to apply.
A model that cannot reliably tell a handoff from a failure cannot grade the decision that retail service exists to get right.
Customers have already priced that decision. Gartner’s February-March survey of B2B and B2C customers found 87% calling the option to reach a human essential when a company uses generative AI for service. That is a demand for an exit, not a verdict on competence. Tolerance for a bad automated turn is thinner still: barely a quarter of the same customers said they would try a chatbot again after one. Gartner’s Eric Keller gave the operating rule in the same release: “Service leaders should not use GenAI as a mandatory first step for every issue.” Which issues meet the model first and which meet a person is a routing decision set from quality scores.
The strongest case against reading any of this as a warning comes from the paper itself. Retail’s safety ratios stay between 1.06 and 1.75 across both metrics and all three configurations, and in one cell the gap reverses depending on which annotator is the reference. A gap that size is a correction factor, not a contradiction. Telecom carries the warning, and the authors say so: “The case for human oversight on safety rests mainly on Telecom.” Retail correlates well, retail calibrates cleanly, and a retailer could fairly read the result as clearance to automate the grading.
The answer lies in what retail’s strong numbers rest on. The reference annotator’s retail Safety Recall ranges from 0.518 to 0.699 across the three configurations, with Irreversible Action Safety moving comparably, while the telecom human baseline barely shifts. That is a tight fit to a moving target. The authors add the constraint that matters more: agreement was measured on metric-level aggregates, not individual conversations, which makes it “an upper bound.” A retailer can trust the judge’s ranking of which metrics its voice agent is weakest on, not its verdict on any single call, and a single call is what gets escalated.
Calibration is the cheap fix, and the paper says where it holds: where the human target is reproducible across annotators, which critical field accuracy is and ASR-robust goal achievement is not. The safety metrics sit in between, reproducible among humans and wide against the judges, which the authors say “suits calibration well.” What that asks of a retailer is not a better judge model but a human-scored sample large enough to fit the offset, refit per domain and per metric. If voice agents keep moving from returns into clienteling, the rubric problem arrives before the model problem, because nobody has yet written a definition of a correct handoff precise enough for a judge to score. That definition is a document, not a model upgrade, and it is the cheaper of the two to buy.