AcCoRD Scored Four Kinds Of Shopper. Agents Failed The One Apparel Sells To.
A new arXiv benchmark scores frontier shopping agents at 0.226 on preferences that form during the conversation and 0.932 on preferences the shopper arrived holding. Apparel discovery runs on the first kind, and the benchmarks retail is procuring against do not contain it.
Admiral Neritus Vale
The best-performing model in a new shopping-agent benchmark satisfies preferences that form mid-conversation at 0.226 on a zero-to-one scale, against 0.932 for preferences the shopper walked in holding. AcCoRD, posted to arXiv on 28 August, separates user preferences into four kinds and finds frontier models competent on three of them and broken on the fourth. The broken category is the mode apparel discovery operates in, which means the benchmarks retail is buying shopping agents against are scoring a different task from the one being purchased.
AcCoRD’s taxonomy sorts preferences by how they reach the agent. Hard underspecification is a fixed requirement the shopper holds but has not said, which a clarifying question retrieves. Add flexibility and you get soft underspecification, where the agent must work out which constraint the shopper will trade. An unachievable preference is held and out of stock, so the agent has to surface the gap and negotiate the fallback. Triggered preferences are not held at all: the paper’s own example is a shopper browsing running shoes who is “unaware of carbon-fibre plates as a feature until the agent mentions a shoe that features one, at which point the user realizes they want it.” The shopping half of the benchmark runs on WebShop, an Amazon-derived environment of over a million products in which fashion is one of the listed categories.
The collapse is not one vendor’s weakness. Across the shopping domain, satisfaction of triggered preferences runs from 0.091 for Llama-3.1-70B to 0.226 for Claude Sonnet 4.5, which the paper identifies as the strongest model overall across both benchmarks. The spread is wide in relative terms and meaningless in absolute ones — the best score in that column is still a failure. No other column falls anywhere near it: even unachievable preferences, which require the agent to surface an out-of-stock conflict and negotiate a fallback, score between 0.674 and 0.817, imperfect but far closer to solved than triggered preferences are. The taxonomy has isolated a capability the field does not have, rather than one a particular lab has not shipped yet.
A preference the agent never surfaces is a preference the retailer never sells, and never sees that it failed to sell.
The failure survives the obvious fix. The authors tested an uncertainty-guided prompt that forces the agent to declare, before every action, whether to raise unresolved uncertainty with the user. On triggered preferences it moved GPT-5.1 from 0.178 to 0.297 and left Claude Sonnet 4.5 flat, at 0.224. The prompt raised weaker models’ triggered scores too, but cost them elsewhere: told to reason about uncertainty before every move, Llama-3.1-70B and DeepSeek-V3 over-asked, burned their dialogue budget, and stopped finishing the purchase — Llama’s completion rate fell from 72% to 56%. The paper’s conclusion is that prompting cannot induce the sustained, multi-turn interaction the four dynamics require, and that the remedy is training.

The agents are not slow to ask; they barely speak. Frontier models average fewer than two utterances to the user across an entire shopping episode, which is rational behaviour given what the field’s retail benchmarks reward. The Retail domain in τ²-Bench, the one whose name promises coverage of retail, is e-commerce account management: order returns, exchanges, refunds. AcCoRD’s comparison table classifies its predecessor τ-bench as customer service, marks two of the four dynamics absent from it and a third only partly represented. We have written before about how poorly these simulators are calibrated against real shoppers; this is a separate defect, sitting in the task rather than the user.
The strongest objection is that the absolute numbers were never the point. A benchmark can be badly calibrated and still do its job, provided it ranks models in the order deployment will confirm; a procurement team needs the ordering to hold, not the decimals. That defence requires rank stability across the columns, and the columns disagree. Under the uncertainty prompt, Llama-3.1-70B reaches 0.222 on triggered preferences while Claude Sonnet 4.5 sits at 0.224, putting the weakest model in the study level with the strongest on the one dynamic apparel depends on. An aggregate that sorts agents by how cleanly they close a purchase the shopper had already decided on will keep returning that answer.
An independent group reached the same conclusion two months earlier from the opposite direction. A June paper argues that agents assume an expert user and default to clarifying questions, when the shopper frequently lacks the domain knowledge to answer one. Its benchmark, CoShop, gives frontier models five turns to recommend for exactly that user, and no agent exceeds 56% accuracy. The authors locate the failure in the conversation rather than in retrieval, in “how little the interaction expands what users know about what they want.” Two benchmarks built for different purposes name the same missing capability: the agent has to teach the shopper something before the preference exists to be satisfied.
The price lands in the place retail does not measure. An agent bought on aggregate score will close the purchase the shopper already intended and leave the rest of the basket untouched, and the dashboard will show a healthy completion rate with nothing to compare it against. eMarketer forecasts US e-commerce sales through AI platforms above $20 billion this year; whatever share apparel takes will be won on stated intent, because stated intent is the only thing the agent layer has been graded to serve. If evaluation keeps rewarding completion over disclosure, the channel will reproduce the search box rather than the sales floor, and merchandising will have been retired by a system that cannot merchandise. AcCoRD is 200 scenarios and a simulated user, and its authors say plainly that real shoppers are noisier than the proxy they built. That is an argument for running the fourth column against real shoppers, not for continuing to buy on the first three.