Research & Technology Deep Dive (Vale)
A nautilus grading clerk stamps 'MISS' on a card showing an olive coat while a tally board behind him reads EXACT 4.35 and NEAR-EXACT 25.12.

'Same Coat, But Olive' Is a Query the Relevance Score Cannot Grade

In Walmart's own labelled data, a modifier query such as 'same coat, but olive' has 4.35 exact answers and 25.12 near-exact ones — and Recall@K, the number these systems are ranked on, scores every near-exact as a miss. Retailers are buying visual search against a metric blind to the result most likely to close the sale.

Admiral Neritus Vale

A shopper who photographs a coat and types “same, but olive” has asked a question most catalogues cannot answer exactly. In the candidate pool Walmart labelled for its own fashion modifier queries, the average query has 4.35 exact matches and 25.12 near-exact ones, where near-exact means an item that “completely satisfies the intent behind the query” while differing in some minor specification. Recall@K, the headline number on the benchmarks these systems are compared on, scores every one of those near-exacts as a miss. Retailers are selecting retrieval systems against a metric that is blind to their most common correct answer.

The counts come from GradCIR, posted to arXiv on 21 September by a Walmart team that has since deployed it as the retrieval backbone of the company’s visual search. Its complaint is structural. Existing composed-retrieval methods “treat relevance as binary and train on triplets with a single positive target,” a fair simplification for a research dataset and a poor description of a catalogue the paper says holds hundreds of millions of items. Composed image retrieval is the technical name for the modifier query: an image plus a text instruction, which is what the camera icon in a search bar does the moment a shopper types anything at all. Google Lens alone fields nearly 20 billion visual searches every month, which makes the labelling convention underneath these systems a commercial question.

The scarcity of exact matches is a property of the query, not of the catalogue. A shopper reaches for a modifier because the thing they pictured is not on the page; the modifier is the record of a search that has already failed once. Walmart’s figures follow that logic, with exact matches falling from 6.03 per query on plain similarity search to 4.35 once a modifier is attached, and the near-exact band widening to absorb the difference. The harder the request, the more of the right answer migrates into the band the metric scores as zero.

Binary relevance does not undercount the near-miss; it files the near-miss with the error.

Training on the grades instead of the binary produces a measurable gain, and the gain concentrates where the queries are hardest. On Walmart’s own graded benchmark, the same architecture supervised on four relevance levels rather than two lifts NDCG@10 on modifier queries from 0.8376 to 0.8875, a wider move than the one it makes on plain similarity search. The paper’s explanation is the most useful sentence in it: “Collapsing these four cases into a binary labeling scheme discards the information that makes a retrieval ranking useful.” What binary supervision throws away is the ordering, and ordering is the only thing a results page is.

The clearest evidence that the metric is blind sits in the same paper’s other table. On FashionIQ, the public benchmark supervised composed retrieval is still ranked on, GradCIR posts 0.6703 average recall against 0.6641 for the strongest peer-reviewed baseline it cites. Six-tenths of a point is the entire visible distance between the two best supervised systems in the published literature, which is less resolution than a purchase decision needs. GradCIR names a second problem with the public numbers: CIRR and CIRCO draw their images from NLVR2 and COCO, so a retriever trained on catalogue photography is out of distribution before its first query is scored.

Pinterest arrived at the same conclusion from the other end of the funnel. PinPoint, accepted to CVPR 2026 and written by six Pinterest researchers, annotated 7,635 composed queries with human-verified judgments and found an average of 9.1 correct answers to each. It is the largest human-verified multi-answer benchmark built for composed retrieval to date. Two firms sharing nothing but a camera icon and a very large catalogue independently concluded that the labelling convention, not the model, was the binding constraint.

The objection is not new, which is the part that should worry buyers. CIRCO, presented at ICCV 2023, already found “several false negatives” in CIRR, rebuilt the evaluation with an average of 4.53 ground truths per query, and abandoned Recall@K for mean average precision. Three years on, the leaderboards still report Recall@K. A leaderboard is a coordination device, and the cost of leaving one falls entirely on whoever leaves first.

The strongest case for leaving the metric alone is that its blindness may be harmless. If binary-trained and graded-trained retrievers order the top of the page the same way, Recall@K is a noisy proxy rather than a biased one, and a retailer selecting on it loses precision in the comparison but not the system it wanted. Walmart’s table answers that directly: the gap between binary and graded supervision is widest on modifier queries, the case where exact matches are scarcest and the ranking problem is real. A proxy whose error grows with the difficulty of the query favors whichever system is best at the query a shopper could have typed into a filter.

The labels a retailer needs are already sitting in its own catalogue. Walmart built 3.5 million graded training pairs with no manual annotation, using a vision-language model to grade its own assortment against its own queries, which moves graded evaluation from a research programme to a compute line. If that capability stays inside the handful of retailers large enough to build their own retrieval, and never reaches the vendor scorecard, buyers will keep choosing systems that reward a catalogue holding a duplicate and penalise one holding a good substitute. The near-exact is the result that sells the coat in a colour the shopper had not considered, and writing it into the specification is a sentence in a contract rather than a problem for research.