Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913515621580800 |
|---|---|
| author | Xu, Rengan Yang, Junjie Xu, Yifan Li, Hong Liu, Xing Shankar, Devashish Zhang, Haoci Liu, Meng Li, Boyang Hu, Yuxi Tang, Mingwei Zhang, Zehua Zhang, Tunhou Li, Dai Chen, Sijia Musumeci, Gian-Paolo Zhai, Jiaqi Zhu, Bill Yan, Hong Reddy, Srihari |
| author_facet | Xu, Rengan Yang, Junjie Xu, Yifan Li, Hong Liu, Xing Shankar, Devashish Zhang, Haoci Liu, Meng Li, Boyang Hu, Yuxi Tang, Mingwei Zhang, Zehua Zhang, Tunhou Li, Dai Chen, Sijia Musumeci, Gian-Paolo Zhai, Jiaqi Zhu, Bill Yan, Hong Reddy, Srihari |
| contents | The integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models' dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9x speedup and 22x memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3x speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_15373 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention Xu, Rengan Yang, Junjie Xu, Yifan Li, Hong Liu, Xing Shankar, Devashish Zhang, Haoci Liu, Meng Li, Boyang Hu, Yuxi Tang, Mingwei Zhang, Zehua Zhang, Tunhou Li, Dai Chen, Sijia Musumeci, Gian-Paolo Zhai, Jiaqi Zhu, Bill Yan, Hong Reddy, Srihari Machine Learning Artificial Intelligence Information Retrieval The integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models' dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9x speedup and 22x memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3x speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures. |
| title | Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention |
| topic | Machine Learning Artificial Intelligence Information Retrieval |
| url | https://arxiv.org/abs/2409.15373 |