Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tian, Yuxing, Mo, Fengran, Huang, Zhiqi, Zhang, Weixu, Nie, Jian-Yun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908995582689280
author Tian, Yuxing
Mo, Fengran
Huang, Zhiqi
Zhang, Weixu
Nie, Jian-Yun
author_facet Tian, Yuxing
Mo, Fengran
Huang, Zhiqi
Zhang, Weixu
Nie, Jian-Yun
contents Large Language Models (LLMs) have recently been explored as fine-grained zero-shot re-rankers by leveraging attention signals to estimate document relevance. However, existing methods either aggregate attention signals across all heads or rely on a statically selected subset identified by heuristic rules. This solution can be suboptimal because the informative heads can vary across queries or domains. Moreover, naively combining multiple heads can degrade performance due to redundancy or conflicting ranking signals. In this paper, we propose a query-dependent head selection method, RouteHead, for attention-based re-ranking with LLMs. Specifically, we learn a lightweight router that can map each query to an optimal head set, and relevance scores are computed by aggregating attention signals only from these heads. Since query-to-head optimal labels are unavailable, we first construct pseudo labels via an offline search. The router represents each head with a learnable embedding and represents each query using an embedding extracted from the hidden states of the frozen LLM. Then it is trained on the pseudo labels with a sparsity regularizer. Experiments on diverse benchmarks and multiple LLM backbones show that the proposed method consistently outperforms strong baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_24608
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
Tian, Yuxing
Mo, Fengran
Huang, Zhiqi
Zhang, Weixu
Nie, Jian-Yun
Information Retrieval
Artificial Intelligence
Computation and Language
Large Language Models (LLMs) have recently been explored as fine-grained zero-shot re-rankers by leveraging attention signals to estimate document relevance. However, existing methods either aggregate attention signals across all heads or rely on a statically selected subset identified by heuristic rules. This solution can be suboptimal because the informative heads can vary across queries or domains. Moreover, naively combining multiple heads can degrade performance due to redundancy or conflicting ranking signals. In this paper, we propose a query-dependent head selection method, RouteHead, for attention-based re-ranking with LLMs. Specifically, we learn a lightweight router that can map each query to an optimal head set, and relevance scores are computed by aggregating attention signals only from these heads. Since query-to-head optimal labels are unavailable, we first construct pseudo labels via an offline search. The router represents each head with a learnable embedding and represents each query using an embedding extracted from the hidden states of the frozen LLM. Then it is trained on the pseudo labels with a sparsity regularizer. Experiments on diverse benchmarks and multiple LLM backbones show that the proposed method consistently outperforms strong baselines.
title Learning to Route Queries to Heads for Attention-based Re-ranking with Large Language Models
topic Information Retrieval
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2604.24608