Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Rengan, Yang, Junjie, Xu, Yifan, Li, Hong, Liu, Xing, Shankar, Devashish, Zhang, Haoci, Liu, Meng, Li, Boyang, Hu, Yuxi, Tang, Mingwei, Zhang, Zehua, Zhang, Tunhou, Li, Dai, Chen, Sijia, Musumeci, Gian-Paolo, Zhai, Jiaqi, Zhu, Bill, Yan, Hong, Reddy, Srihari
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913515621580800
author Xu, Rengan
Yang, Junjie
Xu, Yifan
Li, Hong
Liu, Xing
Shankar, Devashish
Zhang, Haoci
Liu, Meng
Li, Boyang
Hu, Yuxi
Tang, Mingwei
Zhang, Zehua
Zhang, Tunhou
Li, Dai
Chen, Sijia
Musumeci, Gian-Paolo
Zhai, Jiaqi
Zhu, Bill
Yan, Hong
Reddy, Srihari
author_facet Xu, Rengan
Yang, Junjie
Xu, Yifan
Li, Hong
Liu, Xing
Shankar, Devashish
Zhang, Haoci
Liu, Meng
Li, Boyang
Hu, Yuxi
Tang, Mingwei
Zhang, Zehua
Zhang, Tunhou
Li, Dai
Chen, Sijia
Musumeci, Gian-Paolo
Zhai, Jiaqi
Zhu, Bill
Yan, Hong
Reddy, Srihari
contents The integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models' dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9x speedup and 22x memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3x speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures.
format Preprint
id arxiv_https___arxiv_org_abs_2409_15373
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention
Xu, Rengan
Yang, Junjie
Xu, Yifan
Li, Hong
Liu, Xing
Shankar, Devashish
Zhang, Haoci
Liu, Meng
Li, Boyang
Hu, Yuxi
Tang, Mingwei
Zhang, Zehua
Zhang, Tunhou
Li, Dai
Chen, Sijia
Musumeci, Gian-Paolo
Zhai, Jiaqi
Zhu, Bill
Yan, Hong
Reddy, Srihari
Machine Learning
Artificial Intelligence
Information Retrieval
The integration of hardware accelerators has significantly advanced the capabilities of modern recommendation systems, enabling the exploration of complex ranking paradigms previously deemed impractical. However, the GPU-based computational costs present substantial challenges. In this paper, we demonstrate our development of an efficiency-driven approach to explore these paradigms, moving beyond traditional reliance on native PyTorch modules. We address the specific challenges posed by ranking models' dependence on categorical features, which vary in length and complicate GPU utilization. We introduce Jagged Feature Interaction Kernels, a novel method designed to extract fine-grained insights from long categorical features through efficient handling of dynamically sized tensors. We further enhance the performance of attention mechanisms by integrating Jagged tensors with Flash Attention. Our novel Jagged Flash Attention achieves up to 9x speedup and 22x memory reduction compared to dense attention. Notably, it also outperforms dense flash attention, with up to 3x speedup and 53% more memory efficiency. In production models, we observe 10% QPS improvement and 18% memory savings, enabling us to scale our recommendation systems with longer features and more complex architectures.
title Enhancing Performance and Scalability of Large-Scale Recommendation Systems with Jagged Flash Attention
topic Machine Learning
Artificial Intelligence
Information Retrieval
url https://arxiv.org/abs/2409.15373