Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916688417521664 |
|---|---|
| author | Siracusa, Marco Hsu, Olivia Soria-Pardos, Victor Randall, Joshua Grasset, Arnaud Biscondi, Eric Joseph, Doug Allen, Randy Kjolstad, Fredrik Planas, Miquel Moretó Armejach, Adrià |
| author_facet | Siracusa, Marco Hsu, Olivia Soria-Pardos, Victor Randall, Joshua Grasset, Arnaud Biscondi, Eric Joseph, Doug Allen, Randy Kjolstad, Fredrik Planas, Miquel Moretó Armejach, Adrià |
| contents | Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by offloading these lookups to specialized access units, Decoupled Access-Execute (DAE) processors achieve 2.6$\times$ higher performance and 6.4$\times$ higher performance/watt than GPUs on end-to-end models. Then, we propose the Ember compiler for automatically generating optimized DAE code from PyTorch and TensorFlow. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_09870 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures Siracusa, Marco Hsu, Olivia Soria-Pardos, Victor Randall, Joshua Grasset, Arnaud Biscondi, Eric Joseph, Doug Allen, Randy Kjolstad, Fredrik Planas, Miquel Moretó Armejach, Adrià Hardware Architecture Machine Learning Programming Languages C.1.2; C.1.3; D.3.4 Irregular embedding lookups are a critical bottleneck in recommender models, sparse large language models, and graph learning models. In this paper, we first demonstrate that, by offloading these lookups to specialized access units, Decoupled Access-Execute (DAE) processors achieve 2.6$\times$ higher performance and 6.4$\times$ higher performance/watt than GPUs on end-to-end models. Then, we propose the Ember compiler for automatically generating optimized DAE code from PyTorch and TensorFlow. Conversely from other DAE compilers, Ember features multiple intermediate representations specifically designed for different optimization levels. In this way, Ember can implement all optimizations to match the performance of hand-written code, unlocking the full potential of DAE architectures at scale. |
| title | Ember: A Compiler for Efficient Embedding Operations on Decoupled Access-Execute Architectures |
| topic | Hardware Architecture Machine Learning Programming Languages C.1.2; C.1.3; D.3.4 |
| url | https://arxiv.org/abs/2504.09870 |