CXLAimPod: CXL Memory is all you need in AI era

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yang, Yiwei, Zheng, Yusheng, Chen, Yiqi, Liang, Zheng, Chu, Kexin, Zhou, Zhe, Quinn, Andi, Zhang, Wei
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918128892510208
author Yang, Yiwei
Zheng, Yusheng
Chen, Yiqi
Liang, Zheng
Chu, Kexin
Zhou, Zhe
Quinn, Andi
Zhang, Wei
author_facet Yang, Yiwei
Zheng, Yusheng
Chen, Yiqi
Liang, Zheng
Chu, Kexin
Zhou, Zhe
Quinn, Andi
Zhang, Wei
contents The proliferation of data-intensive applications, ranging from large language models to key-value stores, increasingly stresses memory systems with mixed read-write access patterns. Traditional half-duplex architectures such as DDR5 are ill-suited for such workloads, suffering bus turnaround penalties that reduce their effective bandwidth under mixed read-write patterns. Compute Express Link (CXL) offers a breakthrough with its full-duplex channels, yet this architectural potential remains untapped as existing software stacks are oblivious to this capability. This paper introduces CXLAimPod, an adaptive scheduling framework designed to bridge this software-hardware gap through system support, including cgroup-based hints for application-aware optimization. Our characterization quantifies the opportunity, revealing that CXL systems achieve 55-61% bandwidth improvement at balanced read-write ratios compared to flat DDR5 performance, demonstrating the benefits of full-duplex architecture. To realize this potential, the CXLAimPod framework integrates multiple scheduling strategies with a cgroup-based hint mechanism to navigate the trade-offs between throughput, latency, and overhead. Implemented efficiently within the Linux kernel via eBPF, CXLAimPod delivers significant performance improvements over default CXL configurations. Evaluation on diverse workloads shows 7.4% average improvement for Redis (with up to 150% for specific sequential patterns), 71.6% improvement for LLM text generation, and 9.1% for vector databases, demon-strating that duplex-aware scheduling can effectively exploit CXL's architectural advantages.
format Preprint
id arxiv_https___arxiv_org_abs_2508_15980
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CXLAimPod: CXL Memory is all you need in AI era
Yang, Yiwei
Zheng, Yusheng
Chen, Yiqi
Liang, Zheng
Chu, Kexin
Zhou, Zhe
Quinn, Andi
Zhang, Wei
Operating Systems
The proliferation of data-intensive applications, ranging from large language models to key-value stores, increasingly stresses memory systems with mixed read-write access patterns. Traditional half-duplex architectures such as DDR5 are ill-suited for such workloads, suffering bus turnaround penalties that reduce their effective bandwidth under mixed read-write patterns. Compute Express Link (CXL) offers a breakthrough with its full-duplex channels, yet this architectural potential remains untapped as existing software stacks are oblivious to this capability. This paper introduces CXLAimPod, an adaptive scheduling framework designed to bridge this software-hardware gap through system support, including cgroup-based hints for application-aware optimization. Our characterization quantifies the opportunity, revealing that CXL systems achieve 55-61% bandwidth improvement at balanced read-write ratios compared to flat DDR5 performance, demonstrating the benefits of full-duplex architecture. To realize this potential, the CXLAimPod framework integrates multiple scheduling strategies with a cgroup-based hint mechanism to navigate the trade-offs between throughput, latency, and overhead. Implemented efficiently within the Linux kernel via eBPF, CXLAimPod delivers significant performance improvements over default CXL configurations. Evaluation on diverse workloads shows 7.4% average improvement for Redis (with up to 150% for specific sequential patterns), 71.6% improvement for LLM text generation, and 9.1% for vector databases, demon-strating that duplex-aware scheduling can effectively exploit CXL's architectural advantages.
title CXLAimPod: CXL Memory is all you need in AI era
topic Operating Systems
url https://arxiv.org/abs/2508.15980