Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruggeri, Giuseppe, Andri, Renzo, Pagliari, Daniele Jahier, Cavigelli, Lukas
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915369521774592
author Ruggeri, Giuseppe
Andri, Renzo
Pagliari, Daniele Jahier
Cavigelli, Lukas
author_facet Ruggeri, Giuseppe
Andri, Renzo
Pagliari, Daniele Jahier
Cavigelli, Lukas
contents Deep Recommender Models (DLRMs) inference is a fundamental AI workload accounting for more than 79% of the total AI workload in Meta's data centers. DLRMs' performance bottleneck is found in the embedding layers, which perform many random memory accesses to retrieve small embedding vectors from tables of various sizes. We propose the design of tailored data flows to speedup embedding look-ups. Namely, we propose four strategies to look up an embedding table effectively on one core, and a framework to automatically map the tables asymmetrically to the multiple cores of a SoC. We assess the effectiveness of our method using the Huawei Ascend AI accelerators, comparing it with the default Ascend compiler, and we perform high-level comparisons with Nvidia A100. Results show a speed-up varying from 1.5x up to 6.5x for real workload distributions, and more than 20x for extremely unbalanced distributions. Furthermore, the method proves to be much more independent of the query distribution than the baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2507_01676
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
Ruggeri, Giuseppe
Andri, Renzo
Pagliari, Daniele Jahier
Cavigelli, Lukas
Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
Information Retrieval
C.4; D.1.3; H.3.3; H.3.4
Deep Recommender Models (DLRMs) inference is a fundamental AI workload accounting for more than 79% of the total AI workload in Meta's data centers. DLRMs' performance bottleneck is found in the embedding layers, which perform many random memory accesses to retrieve small embedding vectors from tables of various sizes. We propose the design of tailored data flows to speedup embedding look-ups. Namely, we propose four strategies to look up an embedding table effectively on one core, and a framework to automatically map the tables asymmetrically to the multiple cores of a SoC. We assess the effectiveness of our method using the Huawei Ascend AI accelerators, comparing it with the default Ascend compiler, and we perform high-level comparisons with Nvidia A100. Results show a speed-up varying from 1.5x up to 6.5x for real workload distributions, and more than 20x for extremely unbalanced distributions. Furthermore, the method proves to be much more independent of the query distribution than the baseline.
title Deep Recommender Models Inference: Automatic Asymmetric Data Flow Optimization
topic Distributed, Parallel, and Cluster Computing
Artificial Intelligence
Hardware Architecture
Information Retrieval
C.4; D.1.3; H.3.3; H.3.4
url https://arxiv.org/abs/2507.01676