Wukong: Towards a Scaling Law for Large-Scale Recommendation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Buyun, Luo, Liang, Chen, Yuxin, Nie, Jade, Liu, Xi, Guo, Daifeng, Zhao, Yanli, Li, Shen, Hao, Yuchen, Yao, Yantao, Lakshminarayanan, Guna, Wen, Ellie Dingqiao, Park, Jongsoo, Naumov, Maxim, Chen, Wenlin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913375930286080
author Zhang, Buyun
Luo, Liang
Chen, Yuxin
Nie, Jade
Liu, Xi
Guo, Daifeng
Zhao, Yanli
Li, Shen
Hao, Yuchen
Yao, Yantao
Lakshminarayanan, Guna
Wen, Ellie Dingqiao
Park, Jongsoo
Naumov, Maxim
Chen, Wenlin
author_facet Zhang, Buyun
Luo, Liang
Chen, Yuxin
Nie, Jade
Liu, Xi
Guo, Daifeng
Zhao, Yanli
Li, Shen
Hao, Yuchen
Yao, Yantao
Lakshminarayanan, Guna
Wen, Ellie Dingqiao
Park, Jongsoo
Naumov, Maxim
Chen, Wenlin
contents Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Wukong: Towards a Scaling Law for Large-Scale Recommendation
Zhang, Buyun
Luo, Liang
Chen, Yuxin
Nie, Jade
Liu, Xi
Guo, Daifeng
Zhao, Yanli
Li, Shen
Hao, Yuchen
Yao, Yantao
Lakshminarayanan, Guna
Wen, Ellie Dingqiao
Park, Jongsoo
Naumov, Maxim
Chen, Wenlin
Machine Learning
Artificial Intelligence
Scaling laws play an instrumental role in the sustainable improvement in model quality. Unfortunately, recommendation models to date do not exhibit such laws similar to those observed in the domain of large language models, due to the inefficiencies of their upscaling mechanisms. This limitation poses significant challenges in adapting these models to increasingly more complex real-world datasets. In this paper, we propose an effective network architecture based purely on stacked factorization machines, and a synergistic upscaling strategy, collectively dubbed Wukong, to establish a scaling law in the domain of recommendation. Wukong's unique design makes it possible to capture diverse, any-order of interactions simply through taller and wider layers. We conducted extensive evaluations on six public datasets, and our results demonstrate that Wukong consistently outperforms state-of-the-art models quality-wise. Further, we assessed Wukong's scalability on an internal, large-scale dataset. The results show that Wukong retains its superiority in quality over state-of-the-art models, while holding the scaling law across two orders of magnitude in model complexity, extending beyond 100 GFLOP/example, where prior arts fall short.
title Wukong: Towards a Scaling Law for Large-Scale Recommendation
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2403.02545