A HPX Communication Benchmark: Distributed FFT using Collectives
Fuente:
arXiv
Saved in:
| Main Authors: | Strack, Alexander, Pflüger, Dirk |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
by: Möllmann, Henrik, et al.
Published: (2026)
by: Möllmann, Henrik, et al.
Published: (2026)
Experiences Porting Distributed Applications to Asynchronous Tasks: A Multidimensional FFT Case-study
by: Strack, Alexander, et al.
Published: (2024)
by: Strack, Alexander, et al.
Published: (2024)
Parallel FFTW on RISC-V: A Comparative Study including OpenMP, MPI, and HPX
by: Strack, Alexander, et al.
Published: (2025)
by: Strack, Alexander, et al.
Published: (2025)
Radiation Hydrodynamics at Scale: Comparing MPI and Asynchronous Many-Task Runtimes with FleCSI
by: Strack, Alexander, et al.
Published: (2026)
by: Strack, Alexander, et al.
Published: (2026)
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
by: Thüring, Tim, et al.
Published: (2026)
by: Thüring, Tim, et al.
Published: (2026)
Is RISC-V Ready for Machine Learning? Portable Gaussian Processes Using Asynchronous Tasks
by: Strack, Alexander, et al.
Published: (2026)
by: Strack, Alexander, et al.
Published: (2026)
From Prompts to Performance: Evaluating LLMs for Task-based Parallel Code Generation
by: Bantel, Linus, et al.
Published: (2026)
by: Bantel, Linus, et al.
Published: (2026)
GPRat: Gaussian Process Regression with Asynchronous Tasks
by: Helmann, Maksim, et al.
Published: (2025)
by: Helmann, Maksim, et al.
Published: (2025)
Asynchronous-Many-Task Systems: Challenges and Opportunities -- Scaling an AMR Astrophysics Code on Exascale machines using Kokkos and HPX
by: Daiß, Gregor, et al.
Published: (2024)
by: Daiß, Gregor, et al.
Published: (2024)
An Initial Evaluation of Distributed Graph Algorithms using NWGraph and HPX
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
Overcoming Latency-bound Limitations of Distributed Graph Algorithms using the HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
by: Mohammadiporshokooh, Karame, et al.
Published: (2026)
DaggerFFT: A Distributed FFT Framework Using Task Scheduling in Julia
by: Anvari, Sana Taghipour, et al.
Published: (2026)
by: Anvari, Sana Taghipour, et al.
Published: (2026)
Understanding the Communication Needs of Asynchronous Many-Task Systems -- A Case Study of HPX+LCI
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos
by: Diehl, Patrick, et al.
Published: (2024)
by: Diehl, Patrick, et al.
Published: (2024)
Closing a Source Complexity Gap between Chapel and HPX
by: Atre, Shreyas, et al.
Published: (2025)
by: Atre, Shreyas, et al.
Published: (2025)
HPX -- An open source C++ Standard Library for Parallelism and Concurrency
by: Heller, Thomas, et al.
Published: (2023)
by: Heller, Thomas, et al.
Published: (2023)
HPX with Spack and Singularity Containers: Evaluating Overheads for HPX/Kokkos using an astrophysics application
by: Diehl, Patrick, et al.
Published: (2024)
by: Diehl, Patrick, et al.
Published: (2024)
Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java
by: Diehl, Patrick, et al.
Published: (2023)
by: Diehl, Patrick, et al.
Published: (2023)
A New Execution Model and Executor for Adaptively Optimizing the Performance of Parallel Algorithms Using HPX Runtime System
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
by: Mohammadiporshokooh, Karame, et al.
Published: (2025)
TurboFNO: High-Performance Fourier Neural Operator with Fused FFT-GEMM-iFFT on GPU
by: Wu, Shixun, et al.
Published: (2025)
by: Wu, Shixun, et al.
Published: (2025)
Exploring Performance-Productivity Trade-offs in AMT Runtimes: A Task Bench Study of Itoyori, ItoyoriFBC, HPX, and MPI
by: Lahnor, Torben R., et al.
Published: (2026)
by: Lahnor, Torben R., et al.
Published: (2026)
TurboFFT: A High-Performance Fast Fourier Transform with Fault Tolerance on GPU
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
A Framework for Hybrid Collective Inference in Distributed Sensor Networks
by: Nash, Andrew, et al.
Published: (2026)
by: Nash, Andrew, et al.
Published: (2026)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
by: Wu, Shixun, et al.
Published: (2024)
by: Wu, Shixun, et al.
Published: (2024)
Application-Centric Benchmarking of Distributed FaaS Platforms using BeFaaS
by: Grambow, Martin, et al.
Published: (2023)
by: Grambow, Martin, et al.
Published: (2023)
CoCoDiff: Optimizing Collective Communications for Distributed Diffusion Transformer Inference Under Ulysses Sequence Parallelism
by: Ma, Bin, et al.
Published: (2026)
by: Ma, Bin, et al.
Published: (2026)
HiCCL: A Hierarchical Collective Communication Library
by: Hidayetoglu, Mert, et al.
Published: (2024)
by: Hidayetoglu, Mert, et al.
Published: (2024)
Exploiting Multicast for Accelerating Collective Communication
by: Xu, Chao, et al.
Published: (2026)
by: Xu, Chao, et al.
Published: (2026)
Prime Collective Communications Library -- Technical Report
by: Keiblinger, Michael, et al.
Published: (2025)
by: Keiblinger, Michael, et al.
Published: (2025)
exaCB: Reproducible Continuous Benchmark Collections at Scale Leveraging an Incremental Approach
by: Badwaik, Jayesh, et al.
Published: (2026)
by: Badwaik, Jayesh, et al.
Published: (2026)
ZCCL: Significantly Improving Collective Communication With Error-Bounded Lossy Compression
by: Huang, Jiajun, et al.
Published: (2025)
by: Huang, Jiajun, et al.
Published: (2025)
gZCCL: Compression-Accelerated Collective Communication Framework for GPU Clusters
by: Huang, Jiajun, et al.
Published: (2023)
by: Huang, Jiajun, et al.
Published: (2023)
Towards Experiment Execution in Support of Community Benchmark Workflows for HPC
by: von Laszewski, Gregor, et al.
Published: (2025)
by: von Laszewski, Gregor, et al.
Published: (2025)
Optimizing Distributed ML Communication with Fused Computation-Collective Operations
by: Punniyamurthy, Kishore, et al.
Published: (2023)
by: Punniyamurthy, Kishore, et al.
Published: (2023)
NetSenseML: Network-Adaptive Compression for Efficient Distributed Machine Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
An Efficient, Reliable and Observable Collective Communication Library in Large-scale GPU Training Clusters
by: Zhang, Mingjun, et al.
Published: (2025)
by: Zhang, Mingjun, et al.
Published: (2025)
Distributed Consensus Network: A Modularized Communication Framework and Reliability Probabilistic Analysis
by: Li, Yuetai, et al.
Published: (2025)
by: Li, Yuetai, et al.
Published: (2025)
A Hybrid Communication Approach for Metadata Exchange in Geo-Distributed Fog Environments
by: Kruber, Marvin, et al.
Published: (2023)
by: Kruber, Marvin, et al.
Published: (2023)
Characterizing Communication Patterns in Distributed Large Language Model Inference
by: Xu, Lang, et al.
Published: (2025)
by: Xu, Lang, et al.
Published: (2025)
Similar Items
-
GPU-Resident Gaussian Process Regression Leveraging Asynchronous Tasks with HPX
by: Möllmann, Henrik, et al.
Published: (2026) -
Experiences Porting Distributed Applications to Asynchronous Tasks: A Multidimensional FFT Case-study
by: Strack, Alexander, et al.
Published: (2024) -
Parallel FFTW on RISC-V: A Comparative Study including OpenMP, MPI, and HPX
by: Strack, Alexander, et al.
Published: (2025) -
Radiation Hydrodynamics at Scale: Comparing MPI and Asynchronous Many-Task Runtimes with FleCSI
by: Strack, Alexander, et al.
Published: (2026) -
Comparing the Performance of Heterogeneous Conjugate Gradient and Cholesky Solvers on Various Hardware Using SYCL
by: Thüring, Tim, et al.
Published: (2026)