FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Ye, Yufei, Guo, Wei, Wang, Hao, Zhu, Hong, Ye, Yuyang, Liu, Yong, Guo, Huifeng, Tang, Ruiming, Lian, Defu, Chen, Enhong
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915446695919616
author Ye, Yufei
Guo, Wei
Wang, Hao
Zhu, Hong
Ye, Yuyang
Liu, Yong
Guo, Huifeng
Tang, Ruiming
Lian, Defu
Chen, Enhong
author_facet Ye, Yufei
Guo, Wei
Wang, Hao
Zhu, Hong
Ye, Yuyang
Liu, Yong
Guo, Huifeng
Tang, Ruiming
Lian, Defu
Chen, Enhong
contents Scaling laws for autoregressive generative recommenders reveal potential for larger, more versatile systems but mean greater latency and training costs. To accelerate training and inference, we investigated the recent generative recommendation models HSTU and FuXi-$α$, identifying two efficiency bottlenecks: the indexing operations in relative temporal attention bias and the computation of the query-key attention map. Additionally, we observed that relative attention bias in self-attention mechanisms can also serve as attention maps. Previous works like Synthesizer have shown that alternative forms of attention maps can achieve similar performance, naturally raising the question of whether some attention maps are redundant. Through empirical experiments, we discovered that using the query-key attention map might degrade the model's performance in recommendation tasks. To address these bottlenecks, we propose a new framework applicable to Transformer-like recommendation models. On one hand, we introduce Functional Relative Attention Bias, which avoids the time-consuming operations of the original relative attention bias, thereby accelerating the process. On the other hand, we remove the query-key attention map from the original self-attention layer and design a new Attention-Free Token Mixer module. Furthermore, by applying this framework to FuXi-$α$, we introduce a new model, FuXi-$β$. Experiments across multiple datasets demonstrate that FuXi-$β$ outperforms previous state-of-the-art models and achieves significant acceleration compared to FuXi-$α$, while also adhering to the scaling law. Notably, FuXi-$β$ shows an improvement of 27% to 47% in the NDCG@10 metric on large-scale industrial datasets compared to FuXi-$α$. Our code is available in a public repository: https://github.com/USTC-StarTeam/FuXi-beta
format Preprint
id arxiv_https___arxiv_org_abs_2508_10615
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
Ye, Yufei
Guo, Wei
Wang, Hao
Zhu, Hong
Ye, Yuyang
Liu, Yong
Guo, Huifeng
Tang, Ruiming
Lian, Defu
Chen, Enhong
Information Retrieval
Scaling laws for autoregressive generative recommenders reveal potential for larger, more versatile systems but mean greater latency and training costs. To accelerate training and inference, we investigated the recent generative recommendation models HSTU and FuXi-$α$, identifying two efficiency bottlenecks: the indexing operations in relative temporal attention bias and the computation of the query-key attention map. Additionally, we observed that relative attention bias in self-attention mechanisms can also serve as attention maps. Previous works like Synthesizer have shown that alternative forms of attention maps can achieve similar performance, naturally raising the question of whether some attention maps are redundant. Through empirical experiments, we discovered that using the query-key attention map might degrade the model's performance in recommendation tasks. To address these bottlenecks, we propose a new framework applicable to Transformer-like recommendation models. On one hand, we introduce Functional Relative Attention Bias, which avoids the time-consuming operations of the original relative attention bias, thereby accelerating the process. On the other hand, we remove the query-key attention map from the original self-attention layer and design a new Attention-Free Token Mixer module. Furthermore, by applying this framework to FuXi-$α$, we introduce a new model, FuXi-$β$. Experiments across multiple datasets demonstrate that FuXi-$β$ outperforms previous state-of-the-art models and achieves significant acceleration compared to FuXi-$α$, while also adhering to the scaling law. Notably, FuXi-$β$ shows an improvement of 27% to 47% in the NDCG@10 metric on large-scale industrial datasets compared to FuXi-$α$. Our code is available in a public repository: https://github.com/USTC-StarTeam/FuXi-beta
title FuXi-β: Towards a Lightweight and Fast Large-Scale Generative Recommendation Model
topic Information Retrieval
url https://arxiv.org/abs/2508.10615