Equip Pre-ranking with Target Attention by Residual Quantization

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yutong, Zhu, Yu, Qiao, Yichen, Guan, Ziyu, Shao, Lv, Liu, Tong, Zheng, Bo
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910252846284800
author Li, Yutong
Zhu, Yu
Qiao, Yichen
Guan, Ziyu
Shao, Lv
Liu, Tong
Zheng, Bo
author_facet Li, Yutong
Zhu, Yu
Qiao, Yichen
Guan, Ziyu
Shao, Lv
Liu, Tong
Zheng, Bo
contents The pre-ranking stage in industrial recommendation systems faces a fundamental conflict between efficiency and effectiveness. While powerful models like Target Attention (TA) excel at capturing complex feature interactions in the ranking stage, their high computational cost makes them infeasible for pre-ranking, which often relies on simplistic vector-product models. This disparity creates a significant performance bottleneck for the entire system. To bridge this gap, we propose TARQ, a novel pre-ranking framework. Inspired by generative models, TARQ's key innovation is to equip pre-ranking with an architecture approximate to TA by Residual Quantization. This allows us to bring the modeling power of TA into the latency-critical pre-ranking stage for the first time, establishing a new state-of-the-art trade-off between accuracy and efficiency. Extensive offline experiments and large-scale online A/B tests at Taobao demonstrate TARQ's significant improvements in ranking performance. Consequently, our model has been fully deployed in production, serving tens of millions of daily active users and yielding substantial business improvements. The code and data are available at https://github.com/zyody/tarq_sigir2026.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Equip Pre-ranking with Target Attention by Residual Quantization
Li, Yutong
Zhu, Yu
Qiao, Yichen
Guan, Ziyu
Shao, Lv
Liu, Tong
Zheng, Bo
Information Retrieval
Artificial Intelligence
Machine Learning
I.2.0; I.5.0; I.7.0
The pre-ranking stage in industrial recommendation systems faces a fundamental conflict between efficiency and effectiveness. While powerful models like Target Attention (TA) excel at capturing complex feature interactions in the ranking stage, their high computational cost makes them infeasible for pre-ranking, which often relies on simplistic vector-product models. This disparity creates a significant performance bottleneck for the entire system. To bridge this gap, we propose TARQ, a novel pre-ranking framework. Inspired by generative models, TARQ's key innovation is to equip pre-ranking with an architecture approximate to TA by Residual Quantization. This allows us to bring the modeling power of TA into the latency-critical pre-ranking stage for the first time, establishing a new state-of-the-art trade-off between accuracy and efficiency. Extensive offline experiments and large-scale online A/B tests at Taobao demonstrate TARQ's significant improvements in ranking performance. Consequently, our model has been fully deployed in production, serving tens of millions of daily active users and yielding substantial business improvements. The code and data are available at https://github.com/zyody/tarq_sigir2026.
title Equip Pre-ranking with Target Attention by Residual Quantization
topic Information Retrieval
Artificial Intelligence
Machine Learning
I.2.0; I.5.0; I.7.0
url https://arxiv.org/abs/2509.16931