UDON: A case for offloading to general purpose compute on CXL memory

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hermes, Jon, Minor, Josh, Wu, Minjun, Patil, Adarsh, Van Hensbergen, Eric
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909159945928704
author Hermes, Jon
Minor, Josh
Wu, Minjun
Patil, Adarsh
Van Hensbergen, Eric
author_facet Hermes, Jon
Minor, Josh
Wu, Minjun
Patil, Adarsh
Van Hensbergen, Eric
contents Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload. We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices. Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead.
format Preprint
id arxiv_https___arxiv_org_abs_2404_02868
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UDON: A case for offloading to general purpose compute on CXL memory
Hermes, Jon
Minor, Josh
Wu, Minjun
Patil, Adarsh
Van Hensbergen, Eric
Emerging Technologies
Upcoming CXL-based disaggregated memory devices feature special purpose units to offload compute to near-memory. In this paper, we explore opportunities for offloading compute to general purpose cores on CXL memory devices, thereby enabling a greater utility and diversity of offload. We study two classes of popular memory intensive applications: ML inference and vector database as candidates for computational offload. The study uses Arm AArch64-based dual-socket NUMA systems to emulate CXL type-2 devices. Our study shows promising results. With our ML inference model partitioning strategy for compute offload, we can place up to 90% data in remote memory with just 20% performance trade-off. Offloading Hierarchical Navigable Small World (HNSW) kernels in vector databases can provide upto 6.87$\times$ performance improvement with under 10% offload overhead.
title UDON: A case for offloading to general purpose compute on CXL memory
topic Emerging Technologies
url https://arxiv.org/abs/2404.02868