Agentic Commerce Deep Dive (Vale)
A nautilus at an inspection desk holding a Russian-language marketplace listing form up to a lamp, while a scoreboard of language columns behind it stays lit except for one dark row marked COMMERCE.

Ten Languages Went Into the Benchmark. Commerce Came Out Worst.

PolyWorkBench is the first agent benchmark to build language variation into the whole execution trajectory rather than the inputs alone. Commerce is the domain where agents break, and the failure mode passes every structural check a retailer would think to run.

Admiral Neritus Vale

Every agentic commerce pitch a retailer heard this year rested on benchmark scores gathered in a single language. PolyWorkBench, posted to arXiv on 7 July by researchers at Beijing Jiaotong University and Tencent’s Weixin AI, is the first to price that assumption across a whole workflow rather than at a single translation step. Of the five workplace domains it tests, commerce is where capable agents collapse.

The collapse is a scoring-regime effect rather than a difficulty effect. Strong all-round models hold 0.85 to 0.95 Grade on knowledge, legal and manufacturing work, which is respectable. Commerce drops those same models to between 0.57 and 0.72. The benchmark’s commerce tasks are cross-border pricing, logistics comparisons, marketplace listings and market launches, where one arithmetic slip or a broken spreadsheet schema voids the whole deliverable. Legal and knowledge tasks appear to award partial credit for structural correctness, so the same error costs a fraction of the run rather than all of it. The overall average therefore flatters agents on exactly the work a retailer most wants to hand over.

The task-level table is what a cross-border operator should read before signing anything. Twenty-seven of the thirty model-and-harness pairs evaluated failed a Russian-to-Vietnamese logistics comparison, the class of routing decision a wholesaler makes weekly. The worst result anywhere in the benchmark is a Russian marketplace listing, where failing agents averaged a Grade of 0.04, a number describing an absent deliverable rather than a poor one. None of these tasks are unusual, and none are what the demo shows you.

None of it arrives as an error message.

The authors name the failure mode, and it should worry anyone running a checkout in more than one language. Alongside comprehension errors they identify cross-lingual coordination errors, where the agent reads the source correctly but cannot keep source and target aligned across a multi-step trajectory, so the final artefact drifts even though every intermediate step looked plausible. Plausible is the operative word. An agent that misreads a Korean invoice throws something a person will notice; one that reads it correctly and then drifts across eleven subsequent steps returns a clean, confident, wrong number.

The benchmark’s own evaluation design shows that this drift survives structural checking. PolyWorkBench scores every run three ways: a task-specific rubric, executable tests, and an LLM judge for coherence and faithfulness. The rubric and the tests agree closely, since both verify the same schemas and the same arithmetic; the judge agrees with neither. Restricted to runs the deterministic evaluators mark as at least half-solved, its correlation with them falls to 0.02, effectively noise, down from 0.23 across the full sample. It is the same divergence between machine-checkable success and judged quality that we put at roughly thirty points in May for shopping agents, now showing up inside one benchmark’s own scoring stack. Whatever monitoring a retailer has wrapped around its agentic checkout is a test suite, not a judge.

Buying the best model does not settle the question, because the scaffolding around it moves the score further than the language does. Claude Opus 4.8 returns 0.923 Pass@1 inside the ClaudeCode harness. Swap the harness and hold the tasks constant: the same model falls to 0.698 inside Codex, a 0.225-point spread across the four harnesses tested — wider than the model’s entire 0.15-point range across all ten languages, from its low in Russian to its high in French, inside ClaudeCode alone. A retailer that names a model in its contract and leaves the harness to the supplier has specified the smaller variable.

A checkout terminal glowing green with a PASSED indicator while the receipt spooling from it prints mismatched currency symbols and garbled mixed-script product names

The strongest objection is that the gap is manufactured by the measurement. In April, Yunsu Kim and colleagues published GAIA-v2-LILT, a re-audit of the GAIA agent benchmark in five non-English languages, arguing that multilingual benchmarks built by machine translation break their own validity through query-answer misalignment and culturally off-target context. Their corrected workflow lifted agent success rates by as much as 32.7% over the minimally translated versions, and their conclusion is blunt: a substantial share of the multilingual performance gap is benchmark-induced measurement error. For the argument here to fail, PolyWorkBench’s commerce dip would have to be the same artefact.

It is not, and the construction pipeline is the reason. PolyWorkBench’s tasks begin from real-world data seeds — receipts, contracts and incident logs — that an author turns into a task specification directly in the target language, which a second author then audits before release. That is a different failure surface from a benchmark translated wholesale out of English, which is the deficiency GAIA-v2-LILT diagnoses. The critique lands on suites like MAPS, which rendered four English agent benchmarks into eleven languages and whose package now carries GAIA-v2-LILT’s data. It does not land here, and its authors concede that substantial gaps remained in many settings even after correction.

The exposure sits in the wrong place, and buyers are already saying so. DeepL’s 2026 Borderless Business report, covered by AI News in April, found only 17% of surveyed international businesses had moved multilingual operations onto large language models or agentic AI, while naming global expansion the largest single driver of language-AI spending. Its respondents sat in the United States, United Kingdom, France, Germany and Japan: firms buying language automation to enter markets, and judging it on evidence from the markets they already hold.

There is a lever in this data, and it is not the one being sold. Mid-tier models recover up to 14 Grade points from best-of-three sampling, because most of their failures are transient rather than structural. The frontier model gains almost nothing from the same technique, so verification compute buys more where the cheaper model runs than where the best one does. If a retailer’s growth is booked in Vietnamese, Russian and Korean while its testing budget sits in its English-language operation, that is an allocation it made rather than a limit it hit. The allocation can be moved, and the revenue it currently protects is not the revenue on the plan.