Generative Recommendation for Large-Scale Advertising
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915907796729856 |
|---|---|
| author | Xue, Ben Liu, Dan Wang, Lixiang Sun, Mingjie Wang, Peng Zhang, Pengfei Shi, Shaoyun Xu, Tianyu Sha, Yunhao Liu, Zhiqiang Kong, Bo Wang, Bo Yang, Hang Xue, Jieting Wang, Junhao Wang, Shengyu Hui, Shuping Ye, Wencai Lin, Xiao Li, Yongzhi Chen, Yuhang Yin, Zhihui Chen, Quan Wen, Shiyang Wu, Wenjin Li, Han Zhou, Guorui Li, Changcheng Jiang, Peng Gai, Kun |
| author_facet | Xue, Ben Liu, Dan Wang, Lixiang Sun, Mingjie Wang, Peng Zhang, Pengfei Shi, Shaoyun Xu, Tianyu Sha, Yunhao Liu, Zhiqiang Kong, Bo Wang, Bo Yang, Hang Xue, Jieting Wang, Junhao Wang, Shengyu Hui, Shuping Ye, Wencai Lin, Xiao Li, Yongzhi Chen, Yuhang Yin, Zhihui Chen, Quan Wen, Shiyang Wu, Wenjin Li, Han Zhou, Guorui Li, Changcheng Jiang, Peng Gai, Kun |
| contents | Generative recommendation has recently attracted widespread attention in industry due to its potential for scaling and stronger model capacity. However, deploying real-time generative recommendation in large-scale advertising requires designs beyond large-language-model (LLM)-style training and serving recipes. We present a production-oriented generative recommender co-designed across architecture, learning, and serving, named GR4AD (Generative Recommendation for ADdvertising). As for tokenization, GR4AD proposes UA-SID (Unified Advertisement Semantic ID) to capture complicated business information. Furthermore, GR4AD introduces LazyAR, a lazy autoregressive decoder that relaxes layer-wise dependencies for short, multi-candidate generation, preserving effectiveness while reducing inference cost, which facilitates scaling under fixed serving budgets. To align optimization with business value, GR4AD employs VSL (Value-Aware Supervised Learning) and proposes RSPO (Ranking-Guided Softmax Preference Optimization), a ranking-aware, list-wise reinforcement learning algorithm that optimizes value-based rewards under list-level metrics for continual online updates. For online inference, we further propose dynamic beam serving, which adapts beam width across generation levels and online load to control compute. Large-scale online A/B tests show up to 4.2% ad revenue improvement over an existing DLRM-based stack, with consistent gains from both model scaling and inference-time scaling. GR4AD has been fully deployed in Kuaishou advertising system with over 400 million users and achieves high-throughput real-time serving. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_22732 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Generative Recommendation for Large-Scale Advertising Xue, Ben Liu, Dan Wang, Lixiang Sun, Mingjie Wang, Peng Zhang, Pengfei Shi, Shaoyun Xu, Tianyu Sha, Yunhao Liu, Zhiqiang Kong, Bo Wang, Bo Yang, Hang Xue, Jieting Wang, Junhao Wang, Shengyu Hui, Shuping Ye, Wencai Lin, Xiao Li, Yongzhi Chen, Yuhang Yin, Zhihui Chen, Quan Wen, Shiyang Wu, Wenjin Li, Han Zhou, Guorui Li, Changcheng Jiang, Peng Gai, Kun Information Retrieval Machine Learning Generative recommendation has recently attracted widespread attention in industry due to its potential for scaling and stronger model capacity. However, deploying real-time generative recommendation in large-scale advertising requires designs beyond large-language-model (LLM)-style training and serving recipes. We present a production-oriented generative recommender co-designed across architecture, learning, and serving, named GR4AD (Generative Recommendation for ADdvertising). As for tokenization, GR4AD proposes UA-SID (Unified Advertisement Semantic ID) to capture complicated business information. Furthermore, GR4AD introduces LazyAR, a lazy autoregressive decoder that relaxes layer-wise dependencies for short, multi-candidate generation, preserving effectiveness while reducing inference cost, which facilitates scaling under fixed serving budgets. To align optimization with business value, GR4AD employs VSL (Value-Aware Supervised Learning) and proposes RSPO (Ranking-Guided Softmax Preference Optimization), a ranking-aware, list-wise reinforcement learning algorithm that optimizes value-based rewards under list-level metrics for continual online updates. For online inference, we further propose dynamic beam serving, which adapts beam width across generation levels and online load to control compute. Large-scale online A/B tests show up to 4.2% ad revenue improvement over an existing DLRM-based stack, with consistent gains from both model scaling and inference-time scaling. GR4AD has been fully deployed in Kuaishou advertising system with over 400 million users and achieves high-throughput real-time serving. |
| title | Generative Recommendation for Large-Scale Advertising |
| topic | Information Retrieval Machine Learning |
| url | https://arxiv.org/abs/2602.22732 |