Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Zheng, Li, Chaofan, Xiao, Shitao, Li, Chaozhuo, Lian, Defu, Shao, Yingxia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913667089432576
author Liu, Zheng
Li, Chaofan
Xiao, Shitao
Li, Chaozhuo
Lian, Defu
Shao, Yingxia
author_facet Liu, Zheng
Li, Chaofan
Xiao, Shitao
Li, Chaozhuo
Lian, Defu
Shao, Yingxia
contents Large language models (LLMs) provide powerful foundations to perform fine-grained text re-ranking. However, they are often prohibitive in reality due to constraints on computation bandwidth. In this work, we propose a \textbf{flexible} architecture called \textbf{Matroyshka Re-Ranker}, which is designed to facilitate \textbf{runtime customization} of model layers and sequence lengths at each layer based on users' configurations. Consequently, the LLM-based re-rankers can be made applicable across various real-world situations. The increased flexibility may come at the cost of precision loss. To address this problem, we introduce a suite of techniques to optimize the performance. First, we propose \textbf{cascaded self-distillation}, where each sub-architecture learns to preserve a precise re-ranking performance from its super components, whose predictions can be exploited as smooth and informative teacher signals. Second, we design a \textbf{factorized compensation mechanism}, where two collaborative Low-Rank Adaptation modules, vertical and horizontal, are jointly employed to compensate for the precision loss resulted from arbitrary combinations of layer and sequence compression. We perform comprehensive experiments based on the passage and document retrieval datasets from MSMARCO, along with all public datasets from BEIR benchmark. In our experiments, Matryoshka Re-Ranker substantially outperforms the existing methods, while effectively preserving its superior performance across various forms of compression and different application scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2501_16302
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width
Liu, Zheng
Li, Chaofan
Xiao, Shitao
Li, Chaozhuo
Lian, Defu
Shao, Yingxia
Computation and Language
Large language models (LLMs) provide powerful foundations to perform fine-grained text re-ranking. However, they are often prohibitive in reality due to constraints on computation bandwidth. In this work, we propose a \textbf{flexible} architecture called \textbf{Matroyshka Re-Ranker}, which is designed to facilitate \textbf{runtime customization} of model layers and sequence lengths at each layer based on users' configurations. Consequently, the LLM-based re-rankers can be made applicable across various real-world situations. The increased flexibility may come at the cost of precision loss. To address this problem, we introduce a suite of techniques to optimize the performance. First, we propose \textbf{cascaded self-distillation}, where each sub-architecture learns to preserve a precise re-ranking performance from its super components, whose predictions can be exploited as smooth and informative teacher signals. Second, we design a \textbf{factorized compensation mechanism}, where two collaborative Low-Rank Adaptation modules, vertical and horizontal, are jointly employed to compensate for the precision loss resulted from arbitrary combinations of layer and sequence compression. We perform comprehensive experiments based on the passage and document retrieval datasets from MSMARCO, along with all public datasets from BEIR benchmark. In our experiments, Matryoshka Re-Ranker substantially outperforms the existing methods, while effectively preserving its superior performance across various forms of compression and different application scenarios.
title Matryoshka Re-Ranker: A Flexible Re-Ranking Architecture With Configurable Depth and Width
topic Computation and Language
url https://arxiv.org/abs/2501.16302