Operations Deep Dive (Vale)
A nautilus examines an agent contract whose RE-TEST line has been left blank, while a three-hour sandglass drains beside it.

Nobody in Procurement Chose the Sample Size

A paper written from inside a production agent puts a number on the cost retail contracts leave out: three hours for one full benchmark run, triggered every time a model, prompt or catalogue moves. What teams run instead is a subset, and its size is set by whoever owns the evaluation budget rather than by whoever signed the service level.

Admiral Neritus Vale

The recurring cost in an agentic retail contract is the grading, not the model call. Retailers buying shopping and service agents have priced inference and integration; evaluation is the line that renews, every time the model, the prompt, the tool schema or the catalogue underneath it changes. Because a full benchmark rerun is unaffordable at that cadence, the agent running in production is graded on a shrinking sample. Nobody in procurement chose its size.

One of the first published cost pictures from inside a production agent arrived on arXiv last Friday, and the figure that matters is a duration. Yining She and Lei Lin, reporting on a deployed analytics agent, write that one pass of its central 519-question benchmark takes about three hours. That is the unit cost of knowing whether the thing still works. The trigger list is theirs: “As a production agent’s models, prompts, tools, and surrounding systems change, developers need to repeatedly evaluate its configurations.” In retail, add the catalogue, the returns policy and the promotional calendar, none of which move on the release schedule of the model underneath.

Retailers do not control most of the clocks that force a re-test. Ask Ralph, the styling assistant Ralph Lauren launched with Microsoft last September, runs on Azure OpenAI and recommends in-stock Polo product inside the brand’s app. Three things underneath it move on three different schedules: the hosted model, the merchandising prompt, and the inventory the answers point at. She and Lin’s agent generated tens of thousands of evaluation runs across development and monitoring, which is what happens when every part has its own release cycle. A retailer that re-tests only when it changes something of its own is testing half the surface.

Cost tracks what the agent is asked to do, and shopping agents are asked to do the expensive thing. Princeton’s Holistic Agent Leaderboard priced agent evaluation across coding, science, web navigation and customer service, and its cheapest benchmark averaged $13 for a single agent’s run. That is the floor of the range, not what a shopping agent looks like. Online Mind2Web, which sends the agent out to navigate live websites, averaged more than $450. A shopping assistant browses a catalogue, calls tools and negotiates with a user who changes their mind, which puts it at the wrong end of that range every time it is re-graded.

What teams do instead of paying is subsample, and the honest ones measure the error that introduces. She and Lin compared random sampling, historical caching, fixed subsets and adaptive testing against several hundred logged runs of the same benchmark. Their best method reproduced the full-run score to within roughly one percentage point while executing well under half the questions. They then declined it. Difficulty-stratified fixed subsets went into production instead, chosen for operational simplicity, and shipped as a menu of subset sizes so the user could trade execution time against fidelity.

The fidelity of the test is a dropdown.

![A shop service counter where a brass dial marked 100, 200, 300, 400 is being turned down to 100, in front of a wall of numbered question cards that are still sealed.](/{{generate: A department-store service counter with a hand-lettered sign reading ‘AGENT RE-TEST’; a clerk’s hand turning a large brass dial whose engraved positions read 100, 200, 300, 400 down to the 100 setting. Behind the counter, a floor-to-ceiling wall of numbered pigeonholes holding question cards, only the bottom row unsealed and the rest still tied with string. A customer waits at the counter with a returns slip. Composition: dial and hand in the foreground right, wall of pigeonholes filling the background, waiting customer small at the left edge. Mood: bureaucratic, quietly ominous.}})

The case against this alarm is strong and it has a literature. Subsampling is a published method with error bars, not a shortcut invented by tired engineers. tinyBenchmarks showed that a hundred curated examples predict a model’s MMLU score to within two points, and item response theory explains why: a well-chosen question carries more information than a pile of redundant ones. If retail agent benchmarks behave like MMLU, the shrinking sample is a saving, and the only error in the contract was ever budgeting for a full run. That is the condition this argument has to survive: the subset must preserve what a retailer needs to know.

It does not, and the paper that made agent subsetting practical says so. Franck Ndzomga’s study of efficient agent benchmarking ran the method across dozens of agent scaffolds and found the asymmetry plainly: rank-order prediction stays stable under distribution shift, while absolute score prediction degrades. His difficulty filter cuts the task count by 44 to 70 percent and keeps the ordering honest, which is the right goal if you are choosing between vendors. A retailer is not choosing between vendors on a Tuesday afternoon. It is asking whether this week’s build still refuses the refund it refused last week, and that is a question about level, not rank.

What costs a retailer money is inconsistency, and inconsistency is the hardest property to subsample for. τ-bench, the tool-and-policy benchmark that introduced the pass^k metric, found leading function-calling agents succeeding on under half its tasks. A mean like that is survivable if it holds still, because a known failure rate can be staffed around. It does not hold still: on repeated attempts at the identical retail task, those agents held up under a quarter of the time. Reliability is measured by repetition, which is the one experiment a shrinking sample cannot afford.

The fix is written down in the same paper that shrank the sample. She and Lin tell developers to recalibrate after material changes to models, prompts, tools or execution systems, and to run the full benchmark periodically to measure the error their subset introduces. That reconciliation run is the line item with no owner: it ships no feature, closes no ticket, and goes first when an agent moves from launch budget to operating budget. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, escalating cost among its reasons, and the costs that do the cancelling arrive after signature. If retailers keep booking evaluation as a launch expense, the measurement they lose first will be the one that would have told them the agent had stopped working. The number of questions behind a service level is a contract term, and it should be priced beside the uptime figure.