The Agent Buying Your Coat Can't Verify the One Selling It
Agentic checkout is shipping into retail on rails that prove a human authorized the purchase but not that the two agents brokering it can trust each other. The first empirical study of ERC-8004, the protocol meant to supply that proof, finds its identity layer almost entirely placeholder and its reputation scores dominated by coordinated fakes.
Neritus Vale
Google’s flagship example of agent-led shopping is a coat: tell your agent you want a winter jacket in green and will pay up to twenty percent more, and it buys the exact variant the moment one appears. For a retailer to allow that, its agent and the shopper’s agent each need proof the other is real, solvent, and not an impostor. The first empirical study of ERC-8004, the protocol built to supply that proof, finds it almost entirely hollow.
Agentic checkout has moved from slide to storefront. ChatGPT shoppers in the United States can now buy from Etsy sellers, and soon from Shopify brands like Glossier and SKIMS, through the Agentic Commerce Protocol that OpenAI and Stripe open-sourced to run it. Google followed with its Agent Payments Protocol, backed by Mastercard, American Express and Global Fashion Group among more than sixty partners. Both standards do one job well: they let a merchant confirm that a real person authorized a specific purchase, using signed mandates and verifiable credentials. Neither confirms the reverse, which is whether the merchant’s own agent is genuine, funded, and not a counterfeit stall wearing a trusted name. That second question is the one a shopper’s agent has to answer before it hands over a cart and a card.
ERC-8004 is the leading attempt to answer it. Proposed as an Ethereum standard and still a draft, it lets agents “discover, choose, and interact with agents across organizational boundaries without pre-existing trust,” precisely the gap that A2A and MCP, the protocols agents use to talk, were never built to close. It records three things on-chain: an identity for each agent, a ledger of reputation feedback, and a validation registry where independent checkers can vouch for an agent’s work. The design is elegant, permissionless, and already deployed across Ethereum, Base and BNB Chain. The question the study asks is narrower and harder: does what the registry records mean anything?
The identity layer is mostly empty. Xihan Xiong and colleagues, in a study posted to arXiv in June, crawled every ERC-8004 registration on Ethereum, Base and BNB Chain through mid-May and tested which ones resolved to a live agent. Three percent of Ethereum registrations exposed a valid file with a working endpoint; the rest were names pointing at nothing. Base, which reported the highest valid-endpoint share of the three, did best and still cleared only 15 percent. The trust layer a retailer would query to vet a counterpart is, on its flagship chain, almost entirely unoccupied.
The reputation layer is worse, because it looks full. Feedback scores are not comparable across agents, few reviews are tied to a verifiable interaction, and a rating can be bought for the cost of a transaction, so the study set out to see who was actually doing the rating. On Ethereum, 73.5 percent of reviewers moved in coordinated Sybil clusters, wallets acting together to manufacture a history. On Base the share reaches 90.6 percent, which means the busiest reputation market is also the most fabricated. On Base and BSC, most rated agents are left with no genuine track record at all after stripping Sybil feedback; Ethereum fares better, with 84.2 percent retaining at least some non-Sybil history. A five-star merchant agent, then, tells you almost nothing about the merchant, and an agent trained to prefer high ratings would steer straight toward the best-financed fraud.
A layer where almost every identity is a placeholder and almost every rating is manufactured cannot price counterparty risk; it can only pretend to.
The strongest objection is that this is what month one always looks like. Every registry starts as a landgrab of squatters and wash-traded reviews; DNS, app stores and TLS certificates all passed through that swamp before curation, staking and gated validators drained it. For the same to happen here, curation would have to harden faster than agentic checkout scales, and the evidence runs the other way. A companion readiness study found early adoption registration-heavy but operationally shallow: most registered agents were never wired for live operation, and feedback activity was highly concentrated across a handful of wallets. The manipulation is not a phase the network will outgrow. It is the reputation design working as built, at a cost near zero.
For a retailer, the trap closes from both sides. Expose a merchant agent through ACP or AP2 and you become the counterpart a stranger’s agent must verify; fail to, and you are either ignored or impersonated by a stall that has borrowed your name and a bought five-star history. Turn the transaction around, and to trust the incoming shopper agent you are leaning on an identity layer that is, today, mostly empty. The mandates OpenAI, Stripe and Google shipped prove a human stood behind a purchase; they say nothing about whether the two agents brokering it can trust each other, and that is the layer still missing. If checkout agents keep scaling on rails whose counterpart-trust layer stays this thin, the exposure does not vanish; it settles onto the retailer’s balance sheet as returns, chargebacks, and a brand worn by whoever registered the name first. The retailers wiring up agent commerce are not waiting for that layer to become real; they are choosing, mostly in silence, either to trust a counterpart they cannot yet verify or to build the proof themselves.