MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kwon, Miryeong, Gouk, Donghyun, Woo, Hyein, Kim, Junhee, Baek, Jinwoo, Nam, Kyungkuk, Ji, Sangyoon, Kim, Jiseon, Bae, Hanyeoreum, Jang, Junhyeok, You, Hyunwoo, Moon, Junseok, Jung, Myoungsoo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917018829062144
author Kwon, Miryeong
Gouk, Donghyun
Woo, Hyein
Kim, Junhee
Baek, Jinwoo
Nam, Kyungkuk
Ji, Sangyoon
Kim, Jiseon
Bae, Hanyeoreum
Jang, Junhyeok
You, Hyunwoo
Moon, Junseok
Jung, Myoungsoo
author_facet Kwon, Miryeong
Gouk, Donghyun
Woo, Hyein
Kim, Junhee
Baek, Jinwoo
Nam, Kyungkuk
Ji, Sangyoon
Kim, Jiseon
Bae, Hanyeoreum
Jang, Junhyeok
You, Hyunwoo
Moon, Junseok
Jung, Myoungsoo
contents MPI implementations commonly rely on explicit memory-copy operations, incurring overhead from redundant data movement and buffer management. This overhead notably impacts HPC workloads involving intensive inter-processor communication. In response, we introduce MPI-over-CXL, a novel MPI communication paradigm leveraging CXL, which provides cache-coherent shared memory across multiple hosts. MPI-over-CXL replaces traditional data-copy methods with direct shared memory access, significantly reducing communication latency and memory bandwidth usage. By mapping shared memory regions directly into the virtual address spaces of MPI processes, our design enables efficient pointer-based communication, eliminating redundant copying operations. To validate this approach, we implement a comprehensive hardware and software environment, including a custom CXL 3.2 controller, FPGA-based multi-host emulation, and dedicated software stack. Our evaluations using representative benchmarks demonstrate substantial performance improvements over conventional MPI systems, underscoring MPI-over-CXL's potential to enhance efficiency and scalability in large-scale HPC environments.
format Preprint
id arxiv_https___arxiv_org_abs_2510_14622
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
Kwon, Miryeong
Gouk, Donghyun
Woo, Hyein
Kim, Junhee
Baek, Jinwoo
Nam, Kyungkuk
Ji, Sangyoon
Kim, Jiseon
Bae, Hanyeoreum
Jang, Junhyeok
You, Hyunwoo
Moon, Junseok
Jung, Myoungsoo
Distributed, Parallel, and Cluster Computing
MPI implementations commonly rely on explicit memory-copy operations, incurring overhead from redundant data movement and buffer management. This overhead notably impacts HPC workloads involving intensive inter-processor communication. In response, we introduce MPI-over-CXL, a novel MPI communication paradigm leveraging CXL, which provides cache-coherent shared memory across multiple hosts. MPI-over-CXL replaces traditional data-copy methods with direct shared memory access, significantly reducing communication latency and memory bandwidth usage. By mapping shared memory regions directly into the virtual address spaces of MPI processes, our design enables efficient pointer-based communication, eliminating redundant copying operations. To validate this approach, we implement a comprehensive hardware and software environment, including a custom CXL 3.2 controller, FPGA-based multi-host emulation, and dedicated software stack. Our evaluations using representative benchmarks demonstrate substantial performance improvements over conventional MPI systems, underscoring MPI-over-CXL's potential to enhance efficiency and scalability in large-scale HPC environments.
title MPI-over-CXL: Enhancing Communication Efficiency in Distributed HPC Systems
topic Distributed, Parallel, and Cluster Computing
url https://arxiv.org/abs/2510.14622