UDON: A case for offloading to general purpose compute on CXL memory
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909159945928704 |
|---|---|
| author | Hermes, Jon Minor, Josh Wu, Minjun Patil, Adarsh Van Hensbergen, Eric |
| author_facet | Hermes, Jon Minor, Josh Wu, Minjun Patil, Adarsh Van Hensbergen, Eric |
| contents | Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload.
We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices.
Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_02868 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | UDON: A case for offloading to general purpose compute on CXL memory Hermes, Jon Minor, Josh Wu, Minjun Patil, Adarsh Van Hensbergen, Eric Emerging Technologies Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload. We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices. Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead. |
| title | UDON: A case for offloading to general purpose compute on CXL memory |
| topic | Emerging Technologies |
| url | https://arxiv.org/abs/2404.02868 |