The Agent Dropped 37% of Its Tools. Success Stayed Comparable.
A paper posted to arXiv this week formalises when an agent should stop acquiring tools, cutting tool exposure by 37% while task success stayed comparable to full access. Fashion's catalogue work is aimed at ranking; the stopping rule decides how far down the ranking anyone reads.
Sir John Crabstone
A paper posted to arXiv on Wednesday sets a limit on how many tools an agent should acquire. Not which ones. How many, before the next one stops paying for itself. Every brand that shipped an agent endpoint this year is competing for a place above that line.
Scores Are Not Decisions reads a ranked list of candidate tools from the top and stops where the next one no longer earns its cost. Relevance ranking alone, its authors prove, is the wrong rule once tools cost different amounts. The method was evaluated on 1,343 tasks in five tool-use domains, and took the highest payoff of any deployable approach on the retail slice of τ-bench. In live execution it exposed the agent to 37% fewer tools than full access, task success comparable rather than identical. More than a third of the toolkit was surplus.
The harness builders arrived here first. Anthropic’s tool search documentation tells developers to stop loading every tool definition up front. Selection accuracy “degrades once you exceed 30–50 available tools”. The model searches a catalogue instead and loads “only the 3–5 tools” a request needs. A search returns up to five names by default. Only the three to five tools a developer keeps outside the search pool are guaranteed a place; everything else competes for it.
Being reachable is the brand’s decision; being worth the call is not.
Merchants spent this year solving a constraint the platforms had already lifted. Shopify’s Spring ‘26 Edition ships Catalog, Cart and Checkout MCPs, taking a merchant from discovery to purchase inside an agent. Shopify reports that data it syndicates converts at twice the rate in AI chats. Universal availability is the quickest way yet found to make availability worthless.
The advice reaching brands treats this as a hygiene problem. Mapp’s guide to agent-ready catalogues calls schema “the interface layer between your catalogue and the agents”. Its discovery index found that of 400 major fashion brands across seven markets, only 57 had any ChatGPT citations. That counts visibility and is quiet about admission. Clean attributes move a brand up the ranking; they do not decide where the agent stops reading.
Search retail knew what to do about a ranking: it bought position. A stopping rule has no bid box, and no price at which a brand can move itself up the list. That earlier paper’s framework counts context load and privacy exposure against a tool too, not only the cost of the call itself. An integration that asks for more of either is expensive before anyone finds it useful.
What survives a marginal test is what cannot be had cheaper elsewhere. Live stock and a returns rule the agent must honour sit with the brand; the aggregators approximate both. Product copy is the one asset an aggregator can copy outright — the easiest thing to imitate is the thing every brand already publishes for free.
The refusal leaves no trace. Nothing in the logs separates a brand the agent could not find from one it considered and declined. Thin traffic will be filed as weak agent demand, and the money will move to fix the wrong thing.