Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hou, Bojian, Liu, Xiaolong, Liu, Xiaoyi, Xu, Jiaqi, Badr, Yasmine, Hang, Mengyue, Chanpuriya, Sudhanshu, Zhou, Junqing, Yang, Yuhang, Xu, Han, Suo, Qiuling, Chen, Laming, Hu, Yuxi, Zhang, Jiasheng, Xiong, Huaqing, Huang, Yuzhen, Chen, Chao, Dong, Yue, Yang, Yi, Chang, Shuo, Gan, Xiaorui, Chen, Wenlin, Kolay, Santanu, Liu, Darren, Nie, Jade, Yang, Chunzhi, Wen, Ellie, Yang, Jiyan, Li, Huayu
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915797692055552
author Hou, Bojian
Liu, Xiaolong
Liu, Xiaoyi
Xu, Jiaqi
Badr, Yasmine
Hang, Mengyue
Chanpuriya, Sudhanshu
Zhou, Junqing
Yang, Yuhang
Xu, Han
Suo, Qiuling
Chen, Laming
Hu, Yuxi
Zhang, Jiasheng
Xiong, Huaqing
Huang, Yuzhen
Chen, Chao
Dong, Yue
Yang, Yi
Chang, Shuo
Gan, Xiaorui
Chen, Wenlin
Kolay, Santanu
Liu, Darren
Nie, Jade
Yang, Chunzhi
Wen, Ellie
Yang, Jiyan
Li, Huayu
author_facet Hou, Bojian
Liu, Xiaolong
Liu, Xiaoyi
Xu, Jiaqi
Badr, Yasmine
Hang, Mengyue
Chanpuriya, Sudhanshu
Zhou, Junqing
Yang, Yuhang
Xu, Han
Suo, Qiuling
Chen, Laming
Hu, Yuxi
Zhang, Jiasheng
Xiong, Huaqing
Huang, Yuzhen
Chen, Chao
Dong, Yue
Yang, Yi
Chang, Shuo
Gan, Xiaorui
Chen, Wenlin
Kolay, Santanu
Liu, Darren
Nie, Jade
Yang, Chunzhi
Wen, Ellie
Yang, Jiyan
Li, Huayu
contents Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such laws are established for large language models, they remain challenging for recommendation systems, especially those processing both user history and context features. We identify poor scaling efficiency as the main barrier to predictable power-law scaling, stemming from inefficient modules with low Model FLOPs Utilization (MFU) and suboptimal resource allocation. We introduce Kunlun, a scalable architecture that systematically improves model efficiency and resource allocation. Our low-level optimizations include Generalized Dot-Product Attention (GDPA), Hierarchical Seed Pooling (HSP), and Sliding Window Attention. Our high-level innovations feature Computation Skip (CompSkip) and Event-level Personalization. These advances increase MFU from 17% to 37% on NVIDIA B200 GPUs and double scaling efficiency over state-of-the-art methods. Kunlun is now deployed in major Meta Ads models, delivering significant production impact.
format Preprint
id arxiv_https___arxiv_org_abs_2602_10016
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
Hou, Bojian
Liu, Xiaolong
Liu, Xiaoyi
Xu, Jiaqi
Badr, Yasmine
Hang, Mengyue
Chanpuriya, Sudhanshu
Zhou, Junqing
Yang, Yuhang
Xu, Han
Suo, Qiuling
Chen, Laming
Hu, Yuxi
Zhang, Jiasheng
Xiong, Huaqing
Huang, Yuzhen
Chen, Chao
Dong, Yue
Yang, Yi
Chang, Shuo
Gan, Xiaorui
Chen, Wenlin
Kolay, Santanu
Liu, Darren
Nie, Jade
Yang, Chunzhi
Wen, Ellie
Yang, Jiyan
Li, Huayu
Information Retrieval
Artificial Intelligence
Deriving predictable scaling laws that govern the relationship between model performance and computational investment is crucial for designing and allocating resources in massive-scale recommendation systems. While such laws are established for large language models, they remain challenging for recommendation systems, especially those processing both user history and context features. We identify poor scaling efficiency as the main barrier to predictable power-law scaling, stemming from inefficient modules with low Model FLOPs Utilization (MFU) and suboptimal resource allocation. We introduce Kunlun, a scalable architecture that systematically improves model efficiency and resource allocation. Our low-level optimizations include Generalized Dot-Product Attention (GDPA), Hierarchical Seed Pooling (HSP), and Sliding Window Attention. Our high-level innovations feature Computation Skip (CompSkip) and Event-level Personalization. These advances increase MFU from 17% to 37% on NVIDIA B200 GPUs and double scaling efficiency over state-of-the-art methods. Kunlun is now deployed in major Meta Ads models, delivering significant production impact.
title Kunlun: Establishing Scaling Laws for Massive-Scale Recommendation Systems through Unified Architecture Design
topic Information Retrieval
Artificial Intelligence
url https://arxiv.org/abs/2602.10016