A Reranker Cannot Pick What Retrieval Left Out
A CIKM 2026 paper finds that some published gains for LLM rerankers come from a test that plants the answer in the list, and that under realistic retrieval no upgrade it tested significantly beat plain collaborative filtering. Before buying a smarter reranker, a retailer should ask what share of eventual purchases its retrieval stage ever shows the model.
Admiral Neritus Vale
A recommendation reranker can only reorder what retrieval hands it, and a paper posted on arXiv late last month measures how rarely that includes the right product. In The Recall Ceiling of LLM Recommendation Reranking, accepted for oral presentation at CIKM in November, Zhaohui Wang finds that realistic retrieval puts 2–19% of relevant items into a 100-item candidate list across eight datasets. Under those conditions none of the upgrades he tested significantly beat a plain collaborative-filtering (CF) baseline on his three main Amazon datasets. A retailer offered “LLM-powered personalisation” should start by asking what share of the right products its retrieval stage ever surfaces.
Some of the published gains for LLM rerankers come from a test that plants the answer. Part of a 2023 study of LLMs as “zero-shot rankers” scored them on lists that mixed the item each user went on to choose with sampled alternatives. Wang likens the method to “testing a search engine after pre-loading the answer into the index”. On his three Amazon datasets it overstated realistic NDCG@10, a standard score for how high the right items sit in a top-ten list, by 92–95%. The format persists: MemRerank, a March paper on personalised reranking for shopping agents, reports accuracy at picking the right product from five candidates, one of them always correct. Such tests can fairly compare rerankers with one another, but Wang notes that the protocol cannot support “an absolute claim about deployment performance, because the quantity it holds fixed is exactly the one that binds in production.”
A reranker scores zero for any shopper whose next purchase never reached its list, however clever the model. Under the paper’s standard CF retrieval, that was true for 92–98% of users on the three Amazon datasets. Wang turns the point into a ceiling: in his main test, a reranker’s expected NDCG can be no higher than the recall of the list it reads. He concedes the idea is intuitive; what the bound adds is a number against which any reported score can be read.
Making the reranker smarter did not move it past CF. Wang tried prompt engineering, language models spanning a 168-fold range in size, supervised rerankers, LoRA fine-tuning, hybrid text-and-behaviour retrieval and LLM-plus-CF fusion. None produced a statistically significant gain on the three main Amazon datasets. With its “thinking” mode switched on, DeepSeek’s V4-Pro, released in April, scored 43% below CF on Amazon Movies. Reasoning chains, Wang writes, “have nothing to operate on when the relevant items are absent”.
Given CF’s own ranks and scores, the language model got better mainly by deferring to them. Its accuracy rose on every dataset as its ordering moved closer to CF’s. No variant beat CF significantly. In Wang’s words, “telling the model what CF believed mostly teaches it to stay put”. Even a blend of the two, tuned with hindsight on the test users, “recovers CF’s performance and no more”. At these recall levels, what the language model brought of its own was a set of assumptions about which items go together, and those assumptions did more harm than good.
The figure a retailer should ask a personalisation vendor for is recall: the share of products customers went on to buy that sat in the list the model was allowed to read.
The list an LLM reads is often shorter than the one retrieval returns, and lengthening it did not help. The zero-shot LLMs in Wang’s tests read only the top 30 of 100 candidates, a context-budget choice he says matches common practice. On Amazon Beauty that cut the share of relevant items within the model’s view from 8.2% to 3.2%. Widening the list from 10 to 200 candidates raised recall, yet the LLM trailed CF at every length. The recall that binds is the recall of the window, and every candidate added to it lengthens a prompt the vendor pays for.
The best public evidence from fashion sits in the same low-recall band. None of Wang’s eight datasets is an apparel catalogue. H&M’s 2022 Kaggle competition, though, asked teams to predict each customer’s purchases for the following week from two years of real H&M transactions. The winning team’s retrieval stage caught 18.3% of the purchases in the data’s final week within its top 100 candidates per customer, as summarised by Kazuki Fujikawa of the 12th-placed team. Even the best of some 3,000 teams left most of the next week’s purchases where no ranker could reach them.
The strongest objection, which Wang raises himself, is that a live recommender knows far more than any benchmark. Production systems pull candidates from many sources, use signals such as dwell time and session context that his rerankers never saw, and learn from online feedback. If those advantages lift recall to 30% or more, which Wang’s protocol rates adequate, the ceiling loosens and a smarter reranker has room to earn its fee. That is a claim about a measurable number, and the burden of producing it falls on whoever sells the reranker. The closest case in the paper does not favour the seller: MovieLens reached 17.4% recall, about where the H&M winners stood, and all five language models Wang ran there still scored below CF.
Wang’s result leaves retailers a choice about where the money goes. Better retrieval raises the ceiling. The only route around it that his framework allows is generative retrieval, in which the model produces the candidates itself, and he leaves that untested. Shopee already runs one such system, UniRec, which we covered in May. Until a vendor can say what share of eventual purchases its model is shown, a bigger reranker budget buys a costlier judge of a shortlist that, in all eight of Wang’s datasets, usually lacked the answer.