LimRank: Less is More for Reasoning-Intensive Information Reranking

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Song, Tingyu, Zhao, Yilun, Zhang, Siyue, Zhao, Chen, Cohan, Arman
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866917045868691456
author Song, Tingyu
Zhao, Yilun
Zhang, Siyue
Zhao, Chen
Cohan, Arman
author_facet Song, Tingyu
Zhao, Yilun
Zhang, Siyue
Zhao, Chen
Cohan, Arman
contents Existing approaches typically rely on large-scale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive. In this work, we demonstrate that modern LLMs can be effectively adapted using only minimal, high-quality supervision. To enable this, we design LIMRANK-SYNTHESIZER, a reusable and open-source pipeline for generating diverse, challenging, and realistic reranking examples. Using this synthetic data, we fine-tune our reranker model, LIMRANK. We evaluate LIMRANK on two challenging benchmarks, i.e., BRIGHT for reasoning-intensive retrieval and FollowIR for instruction-following retrieval. Our experiments demonstrate that LIMRANK achieves competitive performance, while being trained on less than 5% of the data typically used in prior work. Further ablation studies demonstrate the effectiveness of LIMRANK-SYNTHESIZER and the strong generalization capabilities of LIMRANK across downstream tasks, including scientific literature search and retrieval-augmented generation for knowledge-intensive problem solving.
format Preprint
id arxiv_https___arxiv_org_abs_2510_23544
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LimRank: Less is More for Reasoning-Intensive Information Reranking
Song, Tingyu
Zhao, Yilun
Zhang, Siyue
Zhao, Chen
Cohan, Arman
Computation and Language
Information Retrieval
Existing approaches typically rely on large-scale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive. In this work, we demonstrate that modern LLMs can be effectively adapted using only minimal, high-quality supervision. To enable this, we design LIMRANK-SYNTHESIZER, a reusable and open-source pipeline for generating diverse, challenging, and realistic reranking examples. Using this synthetic data, we fine-tune our reranker model, LIMRANK. We evaluate LIMRANK on two challenging benchmarks, i.e., BRIGHT for reasoning-intensive retrieval and FollowIR for instruction-following retrieval. Our experiments demonstrate that LIMRANK achieves competitive performance, while being trained on less than 5% of the data typically used in prior work. Further ablation studies demonstrate the effectiveness of LIMRANK-SYNTHESIZER and the strong generalization capabilities of LIMRANK across downstream tasks, including scientific literature search and retrieval-augmented generation for knowledge-intensive problem solving.
title LimRank: Less is More for Reasoning-Intensive Information Reranking
topic Computation and Language
Information Retrieval
url https://arxiv.org/abs/2510.23544