Agentic Commerce Deep Dive (Vale)
A merchant robot slides a folded price card across a sealed-bid auction counter while a Visa card lies face-up beside it under a magnifying glass, three identical rival bidders queued behind.

Agents Were Handed the Price. The Benchmark Asks If They Can Hold It.

Visa Research built the first dynamic multi-attribute auction benchmark for LLM pricing agents and found the best of eleven frontier models kept under a third of available profit. The failure is not losing auctions — it is winning them cheaply, which turns a pricing agent into a counterparty that can be read.

Neritus Vale

Visa Research has published the competence test for a product its employer already sells. Bazaar, posted to arXiv on 30 July, puts a language model in the seller’s chair and asks whether it can price against adaptive competitors for customers whose preferences it cannot see. The strongest model tested kept 32.1% of the profit a merchant with perfect hindsight would have taken. The agents were not failing to close; the best of them won most auctions they entered. They were failing to hold margin — a different defect, with a different consequence.

What a retailer buys when it hands pricing to an agent depends on which of those two defects it has. Visa’s Intelligent Commerce portfolio lets agents “discover products, make decisions and complete parts of the purchasing journey” across a network of 4.8 billion payment credentials. The rails assume competence; Bazaar is the first serious attempt to grade it. A configuration-plus-price bid is also the native shape of an apparel decision, whether that is a wholesale quote carrying fabric, delivery window and minimum order, or a markdown carrying depth and timing. An agent that misjudges those occasionally is a cost of doing business. Systematic misjudgement in one direction is not a cost; it is a subsidy, and the party across the counter collects it.

Bazaar’s construction is what makes its verdict usable. Twenty-four synthetic customers hold hidden utility curves over three product attributes, and each round the focal merchant submits one sealed bid, a configuration plus a price, against three specialist bots that offer only their own specialty and compete on price alone. The merchant observes the winner, the winning configuration and the winning price, and nothing else: not customer values, not competitor costs, not losing bids. Because the utilities are closed-form, the authors can compute exactly what full hindsight would have earned, which turns the headline number into a ceiling comparison rather than an impression. In rounds 31 to 40 of the eighty-round game, half the customers silently swap the weight they place on two attributes, and the merchant is given no indication that anything has changed.

The result that matters is not the leaderboard but what predicts it. Across the field, total profit correlates with margin per win at r=0.99 and with win rate at only r=0.88. The skill under test is restraint rather than aggression, and ranking agents by how often they close ranks them by the wrong variable. It happens to be the variable most retail dashboards were built to display.

The inversion is visible at the top of the table. Gemini 3.1 Pro won 77.9% of its auctions to Claude Opus 4.6’s 67.8% and still finished behind it on profit. The authors sort the field’s failures into two named types: Losers, who forfeit most of their regret on auctions the oracle would have won, and Underpricers, who “win most auctions they should but consistently leave surplus on the table.” Underpricing is the more instructive failure, because it reads as success on every metric a merchandising team already tracks.

An underpricer is not a bad negotiator so much as a legible one.

A nautilus in a merchant's apron studies a slate leaderboard{{generate: A nautilus wearing a merchant’s apron, standing at a chalked slate leaderboard in a market stall. The column headed “WINS” is crowded with tally marks; the column headed “MARGIN” is nearly empty. The nautilus holds a piece of chalk but has not written anything. Warm stall light, quiet concentration.}}

Legibility is what converts a pricing weakness into a competitor’s asset. Bazaar’s second finding is that the agents which learn fastest are the ones that revise worst: GPT-5.3 and Opus 4.5 gained 46 and 40 percentage points of win rate before the shock, then ranked among the worst adapters after it, losing an average of nineteen points and never recovering. Reasoning traces show they registered the losses within a few rounds; what they could not do was abandon the hypothesis that had been working. “Revision, not detection, is the bottleneck in post-shock adaptation,” the authors write. A counterparty that converges fast, commits hard and defends a stale belief can be described in one sentence — and anything describable in one sentence can be modelled by the system across the table.

The same model weights produced two different merchants depending on how much thinking they were allowed. GPT-5.4 with reasoning disabled averaged $266 across ten seeds, while the same model at high effort averaged $1,936. More compute moved it from Loser to Underpricer, which is to say it learned to win before it learned to charge, and that ordering should worry anyone signing a contract. A buyer of agentic pricing is procuring a configuration rather than a model, and the vendor’s benchmark card will report the good one.

The strongest case against reading Bazaar this way is that Bazaar is not a market. One language model faces three scripted bots, the customers are synthetic, and the disturbance is a single engineered shock rather than the continuous drift of a real season; the authors list each of these as a limitation. For the argument here to fail, real competitors would have to be noisy enough that a merchant agent’s underpricing signature never resolves into a pattern worth exploiting. That condition does not hold, and the benchmark’s own design shows why. The bots were built to raise margin faster than they cede it, precisely so that “an agent that simply underbids once” could not farm them. An opponent hardened against trivial exploitation is a stiffer test than a live price-follower, not a softer one, and the agents still cleared under a third of what was available.

The public argument about AI and pricing in this industry has been running in the opposite direction. Writing in The Conversation, RMIT’s Aayushi Badhwar warned that shoppers who set price thresholds through an agent may be handing retailers their own reservation price, closing a feedback loop between consumer tools and retail algorithms. That concern is sound, and it covers one side of the counter. Bazaar measures the other side and finds the merchant’s agent leaking surplus in the opposite direction, which means both parties are arriving at agentic commerce with a readable tell.

The useful signal here is not the finding but the filing. A payment network put the caveat to its own product category into the public record before its merchants had reason to look for it, and the caveat is specific: the best agent tested left more than two-thirds of the available profit uncollected. Retailers wiring agents into pricing can read that as a reason to wait or as a specification. The specification is the more useful reading. Instrument margin per win rather than close rate, test the agent against a shift it was never told about, and price on the assumption that whoever sits across the counter is running the same test on you.