Efficient Listwise Reranking with Compressed Document Representations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Déjean, Hervé, Clinchant, Stéphane
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918474120429568
author Déjean, Hervé
Clinchant, Stéphane
author_facet Déjean, Hervé
Clinchant, Stéphane
contents Reranking, the process of refining the output from a first-stage retriever, is often considered computationally expensive, especially when using Large Language Models (LLMs). A common approach to mitigate this cost involves utilizing smaller LLMs or controlling input length. Inspired by recent advances in document compression for retrieval-augmented generation (RAG), we introduce RRK, an efficient and effective listwise reranker compressing documents into multi-token fixed-size embedding representations. Our simple training via distillation shows that this combination of rich compressed representations and listwise reranking yields a highly efficient and effective system. In particular, our 8B-parameter model runs 3x-18x faster than smaller rerankers (0.6-4B parameters) while matching or outperforming them in effectiveness. The efficiency gains are even more striking on long-document benchmarks, where RRK widens its advantage further.
format Preprint
id arxiv_https___arxiv_org_abs_2604_26483
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Efficient Listwise Reranking with Compressed Document Representations
Déjean, Hervé
Clinchant, Stéphane
Information Retrieval
Reranking, the process of refining the output from a first-stage retriever, is often considered computationally expensive, especially when using Large Language Models (LLMs). A common approach to mitigate this cost involves utilizing smaller LLMs or controlling input length. Inspired by recent advances in document compression for retrieval-augmented generation (RAG), we introduce RRK, an efficient and effective listwise reranker compressing documents into multi-token fixed-size embedding representations. Our simple training via distillation shows that this combination of rich compressed representations and listwise reranking yields a highly efficient and effective system. In particular, our 8B-parameter model runs 3x-18x faster than smaller rerankers (0.6-4B parameters) while matching or outperforming them in effectiveness. The efficiency gains are even more striking on long-document benchmarks, where RRK widens its advantage further.
title Efficient Listwise Reranking with Compressed Document Representations
topic Information Retrieval
url https://arxiv.org/abs/2604.26483