LRanker: LLM Ranker for Massive Candidates

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Tao, Lei, Zijie, Hua, Zhigang, Xie, Yan, Yang, Shuang, Liu, Ge, You, Jiaxuan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913166320992256
author Feng, Tao
Lei, Zijie
Hua, Zhigang
Xie, Yan
Yang, Shuang
Liu, Ge
You, Jiaxuan
author_facet Feng, Tao
Lei, Zijie
Hua, Zhigang
Xie, Yan
Yang, Shuang
Liu, Ge
You, Jiaxuan
contents Large language models (LLMs) have recently shown strong potential for ranking by capturing semantic relevance and adapting across diverse domains, yet existing methods remain constrained by limited context length and high computational costs, restricting their applicability to real-world scenarios where candidate pools often scale to millions. To address this challenge, we propose LRanker, a framework tailored for large-candidate ranking. LRanker incorporates a candidate aggregation encoder that leverages K-means clustering to explicitly model global candidate information, and a graph-based test-time scaling mechanism that partitions candidates into subsets, generates multiple query embeddings, and integrates them through an ensemble procedure. By aggregating diverse embeddings instead of relying on a single representation, this mechanism enhances robustness and expressiveness, leading to more accurate ranking over massive candidate pools. We evaluate LRanker on seven tasks across three scenarios in RBench with different candidate scales. Experimental results show that LRanker achieves over 30% gains in the RBench-Small scenario, improves by 3-9% in MRR in the RBench-Large scenario, and sustains scalability with 20-30% improvements in the RBench-Ultra scenario with more than 6.8M candidates. Ablation studies further verify the effectiveness of its key components. Together, these findings demonstrate the robustness, scalability, and effectiveness of LRanker for massive-candidate ranking.
format Preprint
id arxiv_https___arxiv_org_abs_2605_27810
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle LRanker: LLM Ranker for Massive Candidates
Feng, Tao
Lei, Zijie
Hua, Zhigang
Xie, Yan
Yang, Shuang
Liu, Ge
You, Jiaxuan
Information Retrieval
Large language models (LLMs) have recently shown strong potential for ranking by capturing semantic relevance and adapting across diverse domains, yet existing methods remain constrained by limited context length and high computational costs, restricting their applicability to real-world scenarios where candidate pools often scale to millions. To address this challenge, we propose LRanker, a framework tailored for large-candidate ranking. LRanker incorporates a candidate aggregation encoder that leverages K-means clustering to explicitly model global candidate information, and a graph-based test-time scaling mechanism that partitions candidates into subsets, generates multiple query embeddings, and integrates them through an ensemble procedure. By aggregating diverse embeddings instead of relying on a single representation, this mechanism enhances robustness and expressiveness, leading to more accurate ranking over massive candidate pools. We evaluate LRanker on seven tasks across three scenarios in RBench with different candidate scales. Experimental results show that LRanker achieves over 30% gains in the RBench-Small scenario, improves by 3-9% in MRR in the RBench-Large scenario, and sustains scalability with 20-30% improvements in the RBench-Ultra scenario with more than 6.8M candidates. Ablation studies further verify the effectiveness of its key components. Together, these findings demonstrate the robustness, scalability, and effectiveness of LRanker for massive-candidate ranking.
title LRanker: LLM Ranker for Massive Candidates
topic Information Retrieval
url https://arxiv.org/abs/2605.27810