SumRank: Aligning Summarization Models for Long-Document Listwise Reranking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Feng, Jincheng, Liu, Wenhan, Dou, Zhicheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914420914913280
author Feng, Jincheng
Liu, Wenhan
Dou, Zhicheng
author_facet Feng, Jincheng
Liu, Wenhan
Dou, Zhicheng
contents Large Language Models (LLMs) have demonstrated superior performance in listwise passage reranking task. However, directly applying them to rank long-form documents introduces both effectiveness and efficiency issues due to the substantially increased context length. To address this challenge, we propose a pointwise summarization model SumRank, aligned with downstream listwise reranking, to compress long-form documents into concise rank-aligned summaries before the final listwise reranking stage. To obtain our summarization model SumRank, we introduce a three-stage training pipeline comprising cold-start Supervised Fine-Tuning (SFT), specialized RL data construction, and rank-driven alignment via Reinforcement Learning. This paradigm aligns the SumRank with downstream ranking objectives to preserve relevance signals. We conduct extensive experiments on five benchmark datasets from the TREC Deep Learning tracks (TREC DL 19-23). Results show that our lightweight SumRank model achieves state-of-the-art (SOTA) ranking performance while significantly improving efficiency by reducing both summarization overhead and reranking complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2603_24204
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SumRank: Aligning Summarization Models for Long-Document Listwise Reranking
Feng, Jincheng
Liu, Wenhan
Dou, Zhicheng
Information Retrieval
Large Language Models (LLMs) have demonstrated superior performance in listwise passage reranking task. However, directly applying them to rank long-form documents introduces both effectiveness and efficiency issues due to the substantially increased context length. To address this challenge, we propose a pointwise summarization model SumRank, aligned with downstream listwise reranking, to compress long-form documents into concise rank-aligned summaries before the final listwise reranking stage. To obtain our summarization model SumRank, we introduce a three-stage training pipeline comprising cold-start Supervised Fine-Tuning (SFT), specialized RL data construction, and rank-driven alignment via Reinforcement Learning. This paradigm aligns the SumRank with downstream ranking objectives to preserve relevance signals. We conduct extensive experiments on five benchmark datasets from the TREC Deep Learning tracks (TREC DL 19-23). Results show that our lightweight SumRank model achieves state-of-the-art (SOTA) ranking performance while significantly improving efficiency by reducing both summarization overhead and reranking complexity.
title SumRank: Aligning Summarization Models for Long-Document Listwise Reranking
topic Information Retrieval
url https://arxiv.org/abs/2603.24204