Agentic Commerce Deep Dive (Vale)
A nautilus stamps a delivery note PAID IN FULL while the opened crate behind it spills master keys and smaller crates prising open their own lids.

The Mandate Covers the Basket. The Agent Also Opens Accounts.

A September preprint names the layer no agentic-commerce contract contains: nothing in the stack retailers signed this year decides whether an agent should have acquired the account, the service or the subordinate agent it came back with. Payment, budget, OAuth and fulfilment all validate the purchase and stop there.

Neritus Vale

Every agentic-commerce control a retailer bought this year validates the same object: a transaction. Payment limits, budget ceilings, OAuth scopes and fulfilment receipts each answer whether a given purchase was permitted and delivered. None answers whether the agent should have acquired what came back.

Google shipped the retail version of that perimeter on 11 January, launching the Universal Commerce Protocol with Shopify, Etsy, Wayfair, Target and Walmart, endorsed by a list running through Macy’s, Zalando, Visa and Mastercard. The payments layer underneath it, AP2, states its accountability claim without ambiguity: the protocol “provides a non-repudiable, cryptographic audit trail for every transaction, aiding in dispute resolution and building confidence for all participants.” Read the noun. AP2’s own agent authorization specification concedes the boundary, noting that “AP2 makes use of this model for the payments use case, but the model could be applied more generally in the future.” Nothing has been written for the other cases yet.

A preprint posted on 13 September names the layer none of those documents contain. Genliang Zhu and Chu Wang, writing from Georgia Tech, Illinois and Accentrust, call it the post-fulfilment activation gap: “Payment, budget, OAuth, mandate, and fulfillment checks can validate transaction conditions without deciding whether a returned resource may become usable authority.” Their worked case is a task permitted to spend twenty synthetic dollars on one isolated read-only compute instance. The merchant delivers exactly that. The instance resolves to an administrator credential with external-network effects. “Every transaction-side predicate can remain true; the activation predicate must be false.”

The provider’s own paperwork cannot settle the question either. Zhu and Wang classified 1,248 field pairs drawn from five sources of provider evidence and found that no single unit supplied a complete activation profile. A retailer wanting to know what authority it has just taken delivery of must assemble the answer from documents that were never written to agree with each other. No purchase order requires that work, and UCP does not ask for it.

An account is a standing relationship, which is the condition none of these controls were built to evaluate. A cart clears at settlement and takes its obligations with it. What an account carries instead is renewal terms, a credential that outlives the task that opened it, and a contract the retailer is now party to whether or not a human read it. On a resale marketplace the account is the storefront; in wholesale it is a credit line.

The third class of acquisition costs nothing, which is why the payment rail never sees it. A2A, the delegation protocol sitting alongside UCP in the same stack, is built so that “agents collaborate without exposing their internal logic, memory, or proprietary tools.” Opacity is the design goal, not a defect. One of Zhu and Wang’s registered unsafe cases has four aliases independently enrolling four free workers under an envelope that permits a single active descendant. No money moves and no mandate is broken. The task simply ends up commanding four times the labour it was authorised to command.

A budget control can only see the acquisitions that cost something.

A second team reached the same perimeter from the other side. Avital Aviv, Parth A. Gandh, Ron Bitton and Asaf Shabtai catalogued 48 threats against AP2 in August and built working demonstrations for every one they rated highest-risk. Their finding sits before the signature rather than after it: the interactions that shape a transaction, including A2A messages and MCP tool calls, fall outside the protocol’s cryptographic protection. Valid mandate signatures, they conclude, do not ensure that a transaction reflects the user’s intent once its pre-authorisation context has been manipulated. Set the two papers side by side and the mandate is sealed in the middle and unguarded at both ends.

The thesis fails if retailers only ever face agents and never run them. That was a defensible picture of the market a year ago: the shopper’s agent assembles a cart, the retailer stays seller of record, no credential crosses the merchant boundary, and the acquisition problem belongs to whoever hosts the agent. Google’s January launch broke the condition on the day it shipped. Its Business Agent went live with Lowe’s, Michael’s, Poshmark and Reebok, customisable today in Merchant Center and due to train on each retailer’s own data in the coming months. A branded agent that calls tools is an agent that receives things back, and what it receives is the retailer’s to govern.

The most developed model-governance regime in American finance examined this question in April and declined it. OCC Bulletin 2026-13, issued jointly with the Federal Reserve and FDIC as SR 26-2, replaced fifteen-year-old model risk guidance and stated the exclusion twice: “Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance.” That guidance sets supervisory expectations for banks and has no application to a fashion retailer. The absence is instructive anyway, because the supervisors with the longest practice at asking what a model is permitted to do decided the agentic form of the question was not yet answerable. Retailers spent the same year signing protocols that assume it already has been.

The reversal every retailer already owns is a money reversal. Zhu and Wang’s safety proofs cover refunds explicitly, and the result is that a refund does not by itself release a capability that has already been activated. A retailer can be made whole on the purchase and remain exposed on the authority, holding a clean ledger and a live credential at the same time. If retail keeps buying controls that sit at the settlement boundary while its own agents keep acquiring accounts, services and subordinates on the far side of it, the first serious loss will not appear in the chargeback file at all. The decision nobody has written down is the only one left worth writing: what an agent is permitted to come back with.