Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Siracusa, Marco, Hsu, Olivia, Soria-Pardos, Victor, Randall, Joshua, Grasset, Arnaud, Biscondi, Eric, Joseph, Doug, Allen, Randy, Kjolstad, Fredrik, Planas, Miquel Moretó, Armejach, Adrià
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916688417521664
author Siracusa, Marco
Hsu, Olivia
Soria-Pardos, Victor
Randall, Joshua
Grasset, Arnaud
Biscondi, Eric
Joseph, Doug
Allen, Randy
Kjolstad, Fredrik
Planas, Miquel Moretó
Armejach, Adrià
author_facet Siracusa, Marco
Hsu, Olivia
Soria-Pardos, Victor
Randall, Joshua
Grasset, Arnaud
Biscondi, Eric
Joseph, Doug
Allen, Randy
Kjolstad, Fredrik
Planas, Miquel Moretó
Armejach, Adrià
contents Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by offloading these lookups to specialized access units, Decoupled Access-Execute (DAE) processors achieve 2.6$\times$ higher performance and 6.4$\times$ higher performance/watt than GPUs on end-to-end models. Then, we propose the Ember compiler for automatically generating optimized DAE code from PyTorch and TensorFlow. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale.
format Preprint
id arxiv_https___arxiv_org_abs_2504_09870
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
Siracusa, Marco
Hsu, Olivia
Soria-Pardos, Victor
Randall, Joshua
Grasset, Arnaud
Biscondi, Eric
Joseph, Doug
Allen, Randy
Kjolstad, Fredrik
Planas, Miquel Moretó
Armejach, Adrià
Hardware Architecture
Machine Learning
Programming Languages
C.1.2; C.1.3; D.3.4
Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by offloading these lookups to specialized access units, Decoupled Access-Execute (DAE) processors achieve 2.6$\times$ higher performance and 6.4$\times$ higher performance/watt than GPUs on end-to-end models. Then, we propose the Ember compiler for automatically generating optimized DAE code from PyTorch and TensorFlow. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale.
title Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
topic Hardware Architecture
Machine Learning
Programming Languages
C.1.2; C.1.3; D.3.4
url https://arxiv.org/abs/2504.09870