MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866917018829062144 |
|---|---|
| author | Kwon, Miryeong Gouk, Donghyun Woo, Hyein Kim, Junhee Baek, Jinwoo Nam, Kyungkuk Ji, Sangyoon Kim, Jiseon Bae, Hanyeoreum Jang, Junhyeok You, Hyunwoo Moon, Junseok Jung, Myoungsoo |
| author_facet | Kwon, Miryeong Gouk, Donghyun Woo, Hyein Kim, Junhee Baek, Jinwoo Nam, Kyungkuk Ji, Sangyoon Kim, Jiseon Bae, Hanyeoreum Jang, Junhyeok You, Hyunwoo Moon, Junseok Jung, Myoungsoo |
| contents | MPI implementations commonly rely on explicit memory-copy operations, incurring overhead from redundant data movement and buffer management. This overhead notably impacts HPC workloads involving intensive inter-processor communication. In response, we introduce MPI-over-CXL, a novel MPI communication paradigm leveraging CXL, which provides cache-coherent shared memory across multiple hosts. MPI-over-CXL replaces traditional data-copy methods with direct shared memory access, significantly reducing communication latency and memory bandwidth usage. By mapping shared memory regions directly into the virtual address spaces of MPI processes, our design enables efficient pointer-based communication, eliminating redundant copying operations. To validate this approach, we implement a comprehensive hardware and software environment, including a custom CXL 3.2 controller, FPGA-based multi-host emulation, and dedicated software stack. Our evaluations using representative benchmarks demonstrate substantial performance improvements over conventional MPI systems, underscoring MPI-over-CXL's potential to enhance efficiency and scalability in large-scale HPC environments. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2510_14622 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems Kwon, Miryeong Gouk, Donghyun Woo, Hyein Kim, Junhee Baek, Jinwoo Nam, Kyungkuk Ji, Sangyoon Kim, Jiseon Bae, Hanyeoreum Jang, Junhyeok You, Hyunwoo Moon, Junseok Jung, Myoungsoo Distributed, Parallel, and Cluster Computing MPI implementations commonly rely on explicit memory-copy operations, incurring overhead from redundant data movement and buffer management. This overhead notably impacts HPC workloads involving intensive inter-processor communication. In response, we introduce MPI-over-CXL, a novel MPI communication paradigm leveraging CXL, which provides cache-coherent shared memory across multiple hosts. MPI-over-CXL replaces traditional data-copy methods with direct shared memory access, significantly reducing communication latency and memory bandwidth usage. By mapping shared memory regions directly into the virtual address spaces of MPI processes, our design enables efficient pointer-based communication, eliminating redundant copying operations. To validate this approach, we implement a comprehensive hardware and software environment, including a custom CXL 3.2 controller, FPGA-based multi-host emulation, and dedicated software stack. Our evaluations using representative benchmarks demonstrate substantial performance improvements over conventional MPI systems, underscoring MPI-over-CXL's potential to enhance efficiency and scalability in large-scale HPC environments. |
| title | MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems |
| topic | Distributed, Parallel, and Cluster Computing |
| url | https://arxiv.org/abs/2510.14622 |