DeGRe: Dense-supervised Generative Reranking for Recommendation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Chaotian, Zhang, Jingyao, Chen, Chenghao, Sang, Zisen, Zhao, Dehai, Cao, Guodong, Wu, Boxi, Cai, Deng, Jia, Jia
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917531576434688
author Song, Chaotian
Zhang, Jingyao
Chen, Chenghao
Sang, Zisen
Zhao, Dehai
Cao, Guodong
Wu, Boxi
Cai, Deng
Jia, Jia
author_facet Song, Chaotian
Zhang, Jingyao
Chen, Chenghao
Sang, Zisen
Zhao, Dehai
Cao, Guodong
Wu, Boxi
Cai, Deng
Jia, Jia
contents In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequences within an exponentially large permutation space. Recent studies have shifted towards end-to-end generative frameworks, which typically leverage list-wise rewards or preference alignment to guide generator training. However, these methods still face two critical issues. First is the heuristic label bias. Existing methods often construct training targets based on simple rules, such as promoting clicked items to the top, while ignoring causal dependencies within the list context. Second is the credit assignment problem. Sparse list-level posterior rewards fail to directly guide intermediate steps in sequence generation, leading to ambiguous optimization directions. To address these issues, we propose DeGRe (Dense-supervised Generative Reranking), a generative reranking framework that bridges the gap between offline exploration and online efficiency through dense supervision. The core of DeGRe lies in its offline-online decoupled design. During the offline phase, we introduce a Lookahead Evaluator based on cumulative regression, which leverages beam search to actively mine high-value lookahead sequences in the unexposed space. During training, we transform the step-wise value estimations from the evaluator into dense supervision signals and distill them into a lightweight Online Generator. This mechanism enables the generator to internalize lookahead planning capabilities, requiring only a single efficient greedy decoding pass during online inference to approximate the global optimum. Experiments demonstrate that DeGRe outperforms baseline models on public benchmarks and industrial datasets. We have successfully deployed DeGRe on Taobao Flash Shopping, significantly improving online recommendations.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25749
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle DeGRe: Dense-supervised Generative Reranking for Recommendation
Song, Chaotian
Zhang, Jingyao
Chen, Chenghao
Sang, Zisen
Zhao, Dehai
Cao, Guodong
Wu, Boxi
Cai, Deng
Jia, Jia
Information Retrieval
Artificial Intelligence
Machine Learning
In multi-stage recommender systems, reranking optimizes overall utility by capturing intra-list contextual dependencies, yet its central challenge lies in exploring optimal sequences within an exponentially large permutation space. Recent studies have shifted towards end-to-end generative frameworks, which typically leverage list-wise rewards or preference alignment to guide generator training. However, these methods still face two critical issues. First is the heuristic label bias. Existing methods often construct training targets based on simple rules, such as promoting clicked items to the top, while ignoring causal dependencies within the list context. Second is the credit assignment problem. Sparse list-level posterior rewards fail to directly guide intermediate steps in sequence generation, leading to ambiguous optimization directions. To address these issues, we propose DeGRe (Dense-supervised Generative Reranking), a generative reranking framework that bridges the gap between offline exploration and online efficiency through dense supervision. The core of DeGRe lies in its offline-online decoupled design. During the offline phase, we introduce a Lookahead Evaluator based on cumulative regression, which leverages beam search to actively mine high-value lookahead sequences in the unexposed space. During training, we transform the step-wise value estimations from the evaluator into dense supervision signals and distill them into a lightweight Online Generator. This mechanism enables the generator to internalize lookahead planning capabilities, requiring only a single efficient greedy decoding pass during online inference to approximate the global optimum. Experiments demonstrate that DeGRe outperforms baseline models on public benchmarks and industrial datasets. We have successfully deployed DeGRe on Taobao Flash Shopping, significantly improving online recommendations.
title DeGRe: Dense-supervised Generative Reranking for Recommendation
topic Information Retrieval
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.25749