Retailers Won Back the Checkout. The Agent Kept the Failures.
A February benchmark finds that the correction a shopper supplies after a bad recommendation teaches a shopping agent more than any clarifying question asked up front. OpenAI's March retreat from checkout left retailers holding the completed orders and the agent operator holding everything that went wrong first.
Neritus Vale
The most useful thing a shopper can do for a shopping agent is fail to buy from it. New benchmark work finds that the abandoned and corrected session teaches an agent more than a clean transaction does, which matters because OpenAI’s March retreat from checkout left retailers holding the clean transactions and the agent operator holding the corrections. Retailers read that retreat as winning back the customer. They won back the half of the session with nothing to learn from.
The paper is Learning Personalized Agents from Human Feedback, published in February by researchers at Meta Superintelligence Labs, Princeton and Duke. It builds an online-shopping benchmark on 20 simulated users, then swaps each persona halfway through to model a customer whose taste has moved. An agent with no memory of past interactions barely moves across the swap, scoring 27.0% after it, which is the control.
The authors then separate the two channels through which an agent can learn from a person: the clarifying question asked before it acts, and the correction supplied after it has already chosen wrong. On the post-swap test the correction channel alone reached 66.9% success, well ahead of the question channel, because a belief an agent holds confidently is one it never thinks to query.
Those users were simulated, which is the limit of the result and the reason it needs a second source. Alibaba’s Learn-to-Ask took the same loop into production, training a proactive agent directly from archived expert conversation logs on the stated grounds that a high-fidelity simulator of open-ended human behaviour is prohibitively hard to build. It was deployed on the company’s medication assistant rather than an apparel catalogue, and that gap is real. What crosses the gap is the mechanism: a pile of ordinary, largely unresolved past conversations was enough to train a policy that converted at 1.87 times the human-staffed service it ran alongside. Nobody in that dataset had been trying to teach anything.
An agent’s failures are not the operating cost of the service; they are the inventory.

OpenAI’s March announcement decided who accumulates that inventory. Merchants may now “use their own checkout experiences while we focus our efforts on product discovery”, as the company put it — a sentence that reads as a concession and works as a division of assets. Target, Sephora, Nordstrom, Best Buy, Lowe’s, The Home Depot and Wayfair are among the retailers integrated for discovery. Discovery is the part where a shopper describes what they want badly, is shown the wrong thing, says no, and tries again. The Agentic Commerce Protocol behind the arrangement publishes two specifications, Agentic Checkout and Delegate Payment, and both describe the transaction. Neither defines a way to return the conversation that preceded it.
The consensus reading of the pivot is that checkout turned out to be hard. Forrester’s Emily Pfeiffer and Sucharita Kodali wrote that OpenAI “made the call to pull the plug on native checkout early”, and the operating numbers back them: Walmart’s Daniel Danker told Wired, as Modern Retail reported, that conversion ran three times lower inside the chatbot than when shoppers clicked out to the site. That explains why OpenAI let go of the transaction. It does not explain why it kept the conversation, and the trade coverage did not ask.
The strongest objection is that these logs are worth less than the argument needs. Real interaction data is not the clean labelled correction a benchmark supplies, and a study of exactly this problem found that harvesting implicit feedback from WildChat and LMSYS conversation logs helped on short designed questions and did not help on longer, more complex ones. Its authors call the signal noisy in their title. For the argument here to fail, that noise would have to be irreducible: unusable at any scale, in any domain. The reason it is reducible in shopping is that shopping resolves. An open-ended chat produces no external verdict on whether the answer was good, while a shopping session ends in a purchase, an abandonment or a return, and each of those labels the transcript that produced it.
What the accumulated failures are eventually used for is the part retailers have not priced. A record of what every catalogue could not answer, pooled across every retailer integrated for discovery, is a product brief no single retailer can assemble from its own orders, because its own orders contain only what worked. If the split holds, with discovery upstream and fulfilment downstream, the party holding the logs will learn which garments should exist before the merchants who make them do. That is the asymmetry marketplace sellers spent a decade complaining about, reappearing one layer above the marketplace. Blocking the agents is no longer available either: on 4 August the Ninth Circuit vacated Amazon’s injunction against Perplexity’s Comet, holding that a user directing an agent is the one accessing the retailer’s computers. Retailers cannot shut the agents out and have not charged them for the logs; only the second of those is still a choice.