Dual-Rerank: Fusing Causality and Utility for Industrial Generative Reranking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Chao, Lin, Shuai, Dai, ChengLei, Qian, Ye, Mingyang, Fan, Zhang, Yi, Wang, Yi, Zhuo, Jingwei
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913016610553856
author Zhang, Chao
Lin, Shuai
Dai, ChengLei
Qian, Ye
Mingyang, Fan
Zhang, Yi
Wang, Yi
Zhuo, Jingwei
author_facet Zhang, Chao
Lin, Shuai
Dai, ChengLei
Qian, Ye
Mingyang, Fan
Zhang, Yi
Wang, Yi
Zhuo, Jingwei
contents Kuaishou serves over 400 million daily active users, processing hundreds of millions of search queries daily against a repository of tens of billions of short videos. As the final decision layer, the reranking stage determines user experience by optimizing whole-page utility. While traditional score-and-sort methods fail to capture combinatorial dependencies, Generative Reranking offers a superior paradigm by directly modeling the permutation probability. However, deploying Generative Reranking in such a high-stakes environment faces a fundamental dual dilemma: 1) the structural trade-off where Autoregressive (AR) models offer superior Sequential modeling but suffer from prohibitive latency, versus Non-Autoregressive (NAR) models that enable efficiency but lack dependency capturing; 2) the optimization gap where Supervised Learning faces challenges in directly optimizing whole-page utility, while Reinforcement Learning (RL) struggles with instability in high-throughput data streams. To resolve this, we propose Dual-Rerank, a unified framework designed for industrial reranking that bridges the structural gap via Sequential Knowledge Distillation and addresses the optimization gap using List-wise Decoupled Reranking Optimization (LDRO) for stable online RL. Extensive A/B testing on production traffic demonstrates that Dual-Rerank achieves State-of-the-Art performance, significantly improving User satisfaction and Watch Time while drastically reducing inference latency compared to AR baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2604_07420
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dual-Rerank: Fusing Causality and Utility for Industrial Generative Reranking
Zhang, Chao
Lin, Shuai
Dai, ChengLei
Qian, Ye
Mingyang, Fan
Zhang, Yi
Wang, Yi
Zhuo, Jingwei
Information Retrieval
Machine Learning
Kuaishou serves over 400 million daily active users, processing hundreds of millions of search queries daily against a repository of tens of billions of short videos. As the final decision layer, the reranking stage determines user experience by optimizing whole-page utility. While traditional score-and-sort methods fail to capture combinatorial dependencies, Generative Reranking offers a superior paradigm by directly modeling the permutation probability. However, deploying Generative Reranking in such a high-stakes environment faces a fundamental dual dilemma: 1) the structural trade-off where Autoregressive (AR) models offer superior Sequential modeling but suffer from prohibitive latency, versus Non-Autoregressive (NAR) models that enable efficiency but lack dependency capturing; 2) the optimization gap where Supervised Learning faces challenges in directly optimizing whole-page utility, while Reinforcement Learning (RL) struggles with instability in high-throughput data streams. To resolve this, we propose Dual-Rerank, a unified framework designed for industrial reranking that bridges the structural gap via Sequential Knowledge Distillation and addresses the optimization gap using List-wise Decoupled Reranking Optimization (LDRO) for stable online RL. Extensive A/B testing on production traffic demonstrates that Dual-Rerank achieves State-of-the-Art performance, significantly improving User satisfaction and Watch Time while drastically reducing inference latency compared to AR baselines.
title Dual-Rerank: Fusing Causality and Utility for Industrial Generative Reranking
topic Information Retrieval
Machine Learning
url https://arxiv.org/abs/2604.07420