Understanding Data Movement in Tightly Coupled Heterogeneous Systems: A Case Study with the Grace Hopper Superchip
Fuente:
arXiv
Saved in:
| Main Authors: | Fusco, Luigi, Khalilov, Mikhail, Chrapek, Marcin, Chukkapalli, Giridhar, Schulthess, Thomas, Hoefler, Torsten |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026)
by: Ahmed, Mahmoud, et al.
Published: (2026)
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024)
by: Khalilov, Mikhail, et al.
Published: (2024)
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
Adventures with Grace Hopper AI Super Chip and the National Research Platform
by: Hurt, J. Alex, et al.
Published: (2024)
by: Hurt, J. Alex, et al.
Published: (2024)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
Automatic BLAS Offloading on Unified Memory Architecture: A Study on NVIDIA Grace-Hopper
by: Li, Junjie, et al.
Published: (2024)
by: Li, Junjie, et al.
Published: (2024)
Cppless: Single-Source and High-Performance Serverless Programming in C++
by: Copik, Marcin, et al.
Published: (2024)
by: Copik, Marcin, et al.
Published: (2024)
SpComm3D: A Framework for Enabling Sparse Communication in 3D Sparse Kernels
by: Abubaker, Nabil, et al.
Published: (2024)
by: Abubaker, Nabil, et al.
Published: (2024)
FaaSKeeper: Learning from Building Serverless Services with ZooKeeper as an Example
by: Copik, Marcin, et al.
Published: (2022)
by: Copik, Marcin, et al.
Published: (2022)
XaaS: Acceleration as a Service to Enable Productive High-Performance Cloud Computing
by: Hoefler, Torsten, et al.
Published: (2024)
by: Hoefler, Torsten, et al.
Published: (2024)
OSMOSIS: Enabling Multi-Tenancy in Datacenter SmartNICs
by: Khalilov, Mikhail, et al.
Published: (2023)
by: Khalilov, Mikhail, et al.
Published: (2023)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
by: Chen, Tiancheng, et al.
Published: (2025)
by: Chen, Tiancheng, et al.
Published: (2025)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
by: Lian, Xinyu, et al.
Published: (2025)
by: Lian, Xinyu, et al.
Published: (2025)
LLAMP: Assessing Network Latency Tolerance of HPC Applications with Linear Programming
by: Shen, Siyuan, et al.
Published: (2024)
by: Shen, Siyuan, et al.
Published: (2024)
XaaS Containers: Performance-Portable Representation With Source and IR Containers
by: Copik, Marcin, et al.
Published: (2025)
by: Copik, Marcin, et al.
Published: (2025)
SeBS-Flow: Benchmarking Serverless Cloud Function Workflows
by: Schmid, Larissa, et al.
Published: (2024)
by: Schmid, Larissa, et al.
Published: (2024)
Resolving Conflicts with Grace: Dynamically Concurrent Universality
by: Kuznetsov, Petr, et al.
Published: (2025)
by: Kuznetsov, Petr, et al.
Published: (2025)
Core Hours and Carbon Credits: Incentivizing Sustainability in HPC
by: Kamatar, Alok, et al.
Published: (2025)
by: Kamatar, Alok, et al.
Published: (2025)
Inductive Loop Analysis for Practical HPC Application Optimization
by: Schaad, Philipp, et al.
Published: (2025)
by: Schaad, Philipp, et al.
Published: (2025)
Alps, a versatile research infrastructure
by: Martinasso, Maxime, et al.
Published: (2025)
by: Martinasso, Maxime, et al.
Published: (2025)
High Performance Unstructured SpMM Computation Using Tensor Cores
by: Okanovic, Patrik, et al.
Published: (2024)
by: Okanovic, Patrik, et al.
Published: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
by: Chen, Chang, et al.
Published: (2025)
by: Chen, Chang, et al.
Published: (2025)
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers
by: Lepori, Andrea, et al.
Published: (2025)
by: Lepori, Andrea, et al.
Published: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025)
by: Shi, Ruimin, et al.
Published: (2025)
ADELIA: Automatic Differentiation for Efficient Laplace Inference Approximations
by: Boudaoud, Afif, et al.
Published: (2026)
by: Boudaoud, Afif, et al.
Published: (2026)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
by: Proaño, Andrès Rubio, et al.
Published: (2024)
by: Proaño, Andrès Rubio, et al.
Published: (2024)
SuperInfer: SLO-Aware Rotary Scheduling and Memory Management for LLM Inference on Superchips
by: Yu, Jiahuan, et al.
Published: (2026)
by: Yu, Jiahuan, et al.
Published: (2026)
Tight Bounds on Channel Reliability via Generalized Quorum Systems (Extended Version)
by: Naser-Pastoriza, Alejandro, et al.
Published: (2025)
by: Naser-Pastoriza, Alejandro, et al.
Published: (2025)
PICO: Performance Insights for Collective Operations
by: Pasqualoni, Saverio, et al.
Published: (2025)
by: Pasqualoni, Saverio, et al.
Published: (2025)
Towards Specialized Supercomputers for Climate Sciences: Computational Requirements of the Icosahedral Nonhydrostatic Weather and Climate Model
by: Hoefler, Torsten, et al.
Published: (2024)
by: Hoefler, Torsten, et al.
Published: (2024)
FPsPIN: An FPGA-based Open-Hardware Research Platform for Processing in the Network
by: Schneider, Timo, et al.
Published: (2024)
by: Schneider, Timo, et al.
Published: (2024)
Reconfigurable Heterogeneous Quorum Systems
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Tight Lower Bounds in the Supported LOCAL Model
by: Balliu, Alkida, et al.
Published: (2024)
by: Balliu, Alkida, et al.
Published: (2024)
WOW: Workflow-Aware Data Movement and Task Scheduling for Dynamic Scientific Workflows
by: Lehmann, Fabian, et al.
Published: (2025)
by: Lehmann, Fabian, et al.
Published: (2025)
Evolving HPC services to enable ML workloads on HPE Cray EX
by: Schuppli, Stefano, et al.
Published: (2025)
by: Schuppli, Stefano, et al.
Published: (2025)
Understanding the Communication Needs of Asynchronous Many-Task Systems -- A Case Study of HPX+LCI
by: Yan, Jiakun, et al.
Published: (2025)
by: Yan, Jiakun, et al.
Published: (2025)
Coordinated Power Management on Heterogeneous Systems
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Tight Conditions for Binary-Output Tasks under Crashes
by: Albouy, Timothé, et al.
Published: (2025)
by: Albouy, Timothé, et al.
Published: (2025)
Similar Items
-
Cross-Layer Energy Analysis of Multimodal Training on Grace Hopper Superchips
by: Ahmed, Mahmoud, et al.
Published: (2026) -
Network-Offloaded Bandwidth-Optimal Broadcast and Allgather for Distributed AI
by: Khalilov, Mikhail, et al.
Published: (2024) -
Software Resource Disaggregation for HPC with Serverless Computing
by: Copik, Marcin, et al.
Published: (2024) -
Hazel: Secure and Efficient Disaggregated Storage
by: Chrapek, Marcin, et al.
Published: (2025) -
Adventures with Grace Hopper AI Super Chip and the National Research Platform
by: Hurt, J. Alex, et al.
Published: (2024)