Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
Fuente:
arXiv
Saved in:
| Main Authors: | Kinkead, Shannon, Wesley, Jackson, Schonbein, Whit, DeBonis, David, Dosanjh, Matthew G. F., Bienz, Amanda |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025)
by: Adams, Michael, et al.
Published: (2025)
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026)
by: Bridges, Patrick G., et al.
Published: (2026)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025)
by: Collom, Gerald, et al.
Published: (2025)
Otus Supercomputer
by: Ehtesabi, Sadaf, et al.
Published: (2025)
by: Ehtesabi, Sadaf, et al.
Published: (2025)
A More Scalable Sparse Dynamic Data Exchange
by: Geyko, Andrew, et al.
Published: (2023)
by: Geyko, Andrew, et al.
Published: (2023)
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
by: Baskakova, Elena, et al.
Published: (2026)
by: Baskakova, Elena, et al.
Published: (2026)
A Digital Twin Framework for Liquid-cooled Supercomputers as Demonstrated at Exascale
by: Brewer, Wesley, et al.
Published: (2024)
by: Brewer, Wesley, et al.
Published: (2024)
Supercomputing for High-speed Avoidance and Reactive Planning in Robots
by: Lachmansingh, Kieran S., et al.
Published: (2025)
by: Lachmansingh, Kieran S., et al.
Published: (2025)
Enabling Message Passing Interface Containers on the LUMI Supercomputer
by: Lazzaro, Alfio
Published: (2024)
by: Lazzaro, Alfio
Published: (2024)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)
by: Adefemi, Temitayo
Published: (2024)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Supercomputer 3D Digital Twin for User Focused Real-Time Monitoring
by: Bergeron, William, et al.
Published: (2024)
by: Bergeron, William, et al.
Published: (2024)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
by: Cornelius, Melanie, et al.
Published: (2025)
by: Cornelius, Melanie, et al.
Published: (2025)
Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
by: Shubham, et al.
Published: (2024)
by: Shubham, et al.
Published: (2024)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
by: Ma, Xiaolong, et al.
Published: (2024)
by: Ma, Xiaolong, et al.
Published: (2024)
Configurable Non-uniform All-to-all Algorithms
by: Fan, Ke, et al.
Published: (2024)
by: Fan, Ke, et al.
Published: (2024)
FlashFuser: Expanding the Scale of Kernel Fusion for Compute-Intensive Operators via Inter-Core Connection
by: Huang, Ziyu, et al.
Published: (2025)
by: Huang, Ziyu, et al.
Published: (2025)
PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway Supercomputer
by: Jiang, Qingcai, et al.
Published: (2024)
by: Jiang, Qingcai, et al.
Published: (2024)
PICO: Accelerating All k-Core Paradigms on GPU
by: Zhao, Chen, et al.
Published: (2024)
by: Zhao, Chen, et al.
Published: (2024)
CORTEX: Large-Scale Brain Simulator Utilizing Indegree Sub-Graph Decomposition on Fugaku Supercomputer
by: Lyu, Tianxiang, et al.
Published: (2024)
by: Lyu, Tianxiang, et al.
Published: (2024)
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
by: Takahashi, Keichi, et al.
Published: (2023)
by: Takahashi, Keichi, et al.
Published: (2023)
Radiation Hydrodynamics at Scale: Comparing MPI and Asynchronous Many-Task Runtimes with FleCSI
by: Strack, Alexander, et al.
Published: (2026)
by: Strack, Alexander, et al.
Published: (2026)
ColonyOS -- A Meta-Operating System for Distributed Computing Across Heterogeneous Platform
by: Kristiansson, Johan
Published: (2024)
by: Kristiansson, Johan
Published: (2024)
Sustainable Supercomputing for AI: GPU Power Capping at HPC Scale
by: Zhao, Dan, et al.
Published: (2024)
by: Zhao, Dan, et al.
Published: (2024)
Co-designing a Programmable RISC-V Accelerator for MPC-based Energy and Thermal Management of Many-Core HPC Processors
by: Ottaviano, Alessandro, et al.
Published: (2025)
by: Ottaviano, Alessandro, et al.
Published: (2025)
Asynchronous-Many-Task Systems: Challenges and Opportunities -- Scaling an AMR Astrophysics Code on Exascale machines using Kokkos and HPX
by: Daiß, Gregor, et al.
Published: (2024)
by: Daiß, Gregor, et al.
Published: (2024)
Leveraging Core and Uncore Frequency Scaling for Power-Efficient Serverless Workflows
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
by: Tzenetopoulos, Achilleas, et al.
Published: (2024)
Undecided State Dynamics with Many Opinions
by: Cooper, Colin, et al.
Published: (2026)
by: Cooper, Colin, et al.
Published: (2026)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
by: Naman, Pranjal, et al.
Published: (2026)
by: Naman, Pranjal, et al.
Published: (2026)
Unfolding an Atomistic World: Atomistic Simulation of Reactor Pressure Vessel Steel Across Year-and-Meter Scales
by: Han, Haozhi, et al.
Published: (2026)
by: Han, Haozhi, et al.
Published: (2026)
Reducing the Impact of I/O Contention in Numerical Weather Prediction Workflows at Scale Using DAOS
by: Manubens, Nicolau, et al.
Published: (2024)
by: Manubens, Nicolau, et al.
Published: (2024)
NET4EXA: Pioneering the Future of Interconnects for Supercomputing and AI
by: Martinelli, Michele, et al.
Published: (2026)
by: Martinelli, Michele, et al.
Published: (2026)
Trace Replay Simulation of MIT SuperCloud for Studying Optimal Sustainability Policies
by: Brewer, Wesley, et al.
Published: (2025)
by: Brewer, Wesley, et al.
Published: (2025)
From Few to Many Faults: Optimal Adaptive Byzantine Agreement
by: Constantinescu, Andrei, et al.
Published: (2025)
by: Constantinescu, Andrei, et al.
Published: (2025)
Fast Algorithms for Scheduling Many-body Correlation Functions on Accelerators
by: Selvitopi, Oguz, et al.
Published: (2025)
by: Selvitopi, Oguz, et al.
Published: (2025)
Contemplating a Lightweight Communication Interface for Asynchronous Many-Task Systems
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Exploring the Emerging Technologies within the Blockchain Landscape
by: Tareq, Mohammad Ali, et al.
Published: (2024)
by: Tareq, Mohammad Ali, et al.
Published: (2024)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
by: Xu, Yuzhuang, et al.
Published: (2026)
by: Xu, Yuzhuang, et al.
Published: (2026)
Understanding Large-Scale Plasma Simulation Challenges for Fusion Energy on Supercomputers
by: Williams, Jeremy J., et al.
Published: (2024)
by: Williams, Jeremy J., et al.
Published: (2024)
Similar Items
-
Optimizing Allreduce Operations for Modern Heterogeneous Architectures with Multiple Processes per GPU
by: Adams, Michael, et al.
Published: (2025) -
Co-Design and Evaluation of a CPU-Free MPI GPU Communication Abstraction and Implementation
by: Bridges, Patrick G., et al.
Published: (2026) -
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026) -
Persistent and Partitioned MPI for Stencil Communication
by: Collom, Gerald, et al.
Published: (2025) -
Otus Supercomputer
by: Ehtesabi, Sadaf, et al.
Published: (2025)