Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866915797692055552 |
|---|---|
| author | Hou, Bojian Liu, Xiaolong Liu, Xiaoyi Xu, Jiaqi Badr, Yasmine Hang, Mengyue Chanpuriya, Sudhanshu Zhou, Junqing Yang, Yuhang Xu, Han Suo, Qiuling Chen, Laming Hu, Yuxi Zhang, Jiasheng Xiong, Huaqing Huang, Yuzhen Chen, Chao Dong, Yue Yang, Yi Chang, Shuo Gan, Xiaorui Chen, Wenlin Kolay, Santanu Liu, Darren Nie, Jade Yang, Chunzhi Wen, Ellie Yang, Jiyan Li, Huayu |
| author_facet | Hou, Bojian Liu, Xiaolong Liu, Xiaoyi Xu, Jiaqi Badr, Yasmine Hang, Mengyue Chanpuriya, Sudhanshu Zhou, Junqing Yang, Yuhang Xu, Han Suo, Qiuling Chen, Laming Hu, Yuxi Zhang, Jiasheng Xiong, Huaqing Huang, Yuzhen Chen, Chao Dong, Yue Yang, Yi Chang, Shuo Gan, Xiaorui Chen, Wenlin Kolay, Santanu Liu, Darren Nie, Jade Yang, Chunzhi Wen, Ellie Yang, Jiyan Li, Huayu |
| contents | Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such laws are established for large language models, they remain challenging for recommendation systems, especially those processing both user history and context features. We identify poor scaling efficiency as the main barrier to predictable power-law scaling, stemming from inefficient modules with low Model FLOPs Utilization (MFU) and suboptimal resource allocation. We introduce Kunlun, a scalable architecture that systematically improves model efficiency and resource allocation. Our low-level optimizations include Generalized Dot-Product Attention (GDPA), Hierarchical Seed Pooling (HSP), and Sliding Window Attention. Our high-level innovations feature Computation Skip (CompSkip) and Event-level Personalization. These advances increase MFU from 17% to 37% on NVIDIA B200 GPUs and double scaling efficiency over state-of-the-art methods. Kunlun is now deployed in major Meta Ads models, delivering significant production impact. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_10016 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design Hou, Bojian Liu, Xiaolong Liu, Xiaoyi Xu, Jiaqi Badr, Yasmine Hang, Mengyue Chanpuriya, Sudhanshu Zhou, Junqing Yang, Yuhang Xu, Han Suo, Qiuling Chen, Laming Hu, Yuxi Zhang, Jiasheng Xiong, Huaqing Huang, Yuzhen Chen, Chao Dong, Yue Yang, Yi Chang, Shuo Gan, Xiaorui Chen, Wenlin Kolay, Santanu Liu, Darren Nie, Jade Yang, Chunzhi Wen, Ellie Yang, Jiyan Li, Huayu Information Retrieval Artificial Intelligence Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such laws are established for large language models, they remain challenging for recommendation systems, especially those processing both user history and context features. We identify poor scaling efficiency as the main barrier to predictable power-law scaling, stemming from inefficient modules with low Model FLOPs Utilization (MFU) and suboptimal resource allocation. We introduce Kunlun, a scalable architecture that systematically improves model efficiency and resource allocation. Our low-level optimizations include Generalized Dot-Product Attention (GDPA), Hierarchical Seed Pooling (HSP), and Sliding Window Attention. Our high-level innovations feature Computation Skip (CompSkip) and Event-level Personalization. These advances increase MFU from 17% to 37% on NVIDIA B200 GPUs and double scaling efficiency over state-of-the-art methods. Kunlun is now deployed in major Meta Ads models, delivering significant production impact. |
| title | Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design |
| topic | Information Retrieval Artificial Intelligence |
| url | https://arxiv.org/abs/2602.10016 |