Understanding Large-Scale HPC System Behavior Through Cluster-Based Visual Analytics
Fuente:
arXiv
Saved in:
| Main Authors: | Austin, Allison, Shilpika, Lam, Yan To Linus, Kuo, Yun-Hsin, Vishwanath, Venkatram, Papka, Michael E., Ma, Kwan-Liu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
An Incremental Multi-Level, Multi-Scale Approach to Assessment of Multifidelity HPC Systems
by: Shilpika, Shilpika, et al.
Published: (2025)
by: Shilpika, Shilpika, et al.
Published: (2025)
Towards Energy Efficient Co-Scheduling in HPC
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
MRSch: Multi-Resource Scheduling for HPC
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
by: Cornelius, Melanie, et al.
Published: (2025)
by: Cornelius, Melanie, et al.
Published: (2025)
Exploring Uncore Frequency Scaling for Heterogeneous Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
Driving Computational Efficiency in Large-Scale Platforms using HPC Technologies
by: Mendez, Alexander Martinez, et al.
Published: (2026)
by: Mendez, Alexander Martinez, et al.
Published: (2026)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
by: Loi, Chung Ming, et al.
Published: (2025)
by: Loi, Chung Ming, et al.
Published: (2025)
DiT-HC: Enabling Efficient Training of Visual Generation Model DiT on HPC-oriented CPU Cluster
by: Zhang, Jinxiao, et al.
Published: (2026)
by: Zhang, Jinxiao, et al.
Published: (2026)
Solutions for Distributed Memory Access Mechanism on HPC Clusters
by: Meizner, Jan, et al.
Published: (2025)
by: Meizner, Jan, et al.
Published: (2025)
Applying Large-Scale Distributed Computing to Structural Bioinformatics -- Bridging Legacy HPC Clusters With Big Data Technologies Using kafka-slurm-agent
by: Rubach, Pawel
Published: (2025)
by: Rubach, Pawel
Published: (2025)
Autonomy Loops for Monitoring, Operational Data Analytics, Feedback, and Response in HPC Operations
by: Boito, Francieli, et al.
Published: (2024)
by: Boito, Francieli, et al.
Published: (2024)
RHAPSODY: Execution of Hybrid AI-HPC Workflows at Scale
by: Alsaadi, Aymen, et al.
Published: (2025)
by: Alsaadi, Aymen, et al.
Published: (2025)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
Resource Optimization with MPI Process Malleability for Dynamic Workloads in HPC Clusters
by: Iserte, Sergio, et al.
Published: (2025)
by: Iserte, Sergio, et al.
Published: (2025)
Distributed Neural Representation for Reactive in situ Visualization
by: Wu, Qi, et al.
Published: (2023)
by: Wu, Qi, et al.
Published: (2023)
More for Less: Integrating Capability-Predominant and Capacity-Predominant Computing
by: Zheng, Zhong, et al.
Published: (2025)
by: Zheng, Zhong, et al.
Published: (2025)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
by: Zheng, Zhong, et al.
Published: (2026)
by: Zheng, Zhong, et al.
Published: (2026)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
by: Zojer, Patrick, et al.
Published: (2026)
by: Zojer, Patrick, et al.
Published: (2026)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
Attack Graph Generation on HPC Clusters
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
Distributed 3D Gaussian Splatting for High-Resolution Isosurface Visualization
by: Han, Mengjiao, et al.
Published: (2025)
by: Han, Mengjiao, et al.
Published: (2025)
Leveraging HPC Profiling & Tracing Tools to Understand the Performance of Particle-in-Cell Monte Carlo Simulations
by: Williams, Jeremy J., et al.
Published: (2023)
by: Williams, Jeremy J., et al.
Published: (2023)
MIDAS: Adaptive Proxy Middleware for Mitigating Metadata Hotspots in HPC I/O at Scale
by: Ghimire, Sangam, et al.
Published: (2025)
by: Ghimire, Sangam, et al.
Published: (2025)
Toward Distributed 3D Gaussian Splatting for High-Resolution Isosurface Visualization
by: Han, Mengjiao, et al.
Published: (2025)
by: Han, Mengjiao, et al.
Published: (2025)
HPC with Enhanced User Separation
by: Prout, Andrew, et al.
Published: (2024)
by: Prout, Andrew, et al.
Published: (2024)
Analysis of the carbon footprint of HPC
by: Benhari, Abdessalam, et al.
Published: (2025)
by: Benhari, Abdessalam, et al.
Published: (2025)
Parallel DNA Sequence Alignment on High-Performance Systems with CUDA and MPI
by: Zwaka, Linus
Published: (2024)
by: Zwaka, Linus
Published: (2024)
Parallel I/O Characterization and Optimization on Large-Scale HPC Systems: A 360-Degree Survey
by: Ather, Hammad, et al.
Published: (2024)
by: Ather, Hammad, et al.
Published: (2024)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
by: Arima, Eishi, et al.
Published: (2024)
by: Arima, Eishi, et al.
Published: (2024)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Understanding the Landscape of Ampere GPU Memory Errors
by: Zhu, Zhu, et al.
Published: (2025)
by: Zhu, Zhu, et al.
Published: (2025)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
by: Ma, Xiaolong, et al.
Published: (2024)
by: Ma, Xiaolong, et al.
Published: (2024)
A Real-Time Digital Twin for Adaptive Scheduling
by: Zhang, Yihe, et al.
Published: (2025)
by: Zhang, Yihe, et al.
Published: (2025)
UNR: Unified Notifiable RMA Library for HPC
by: Feng, Guangnan, et al.
Published: (2024)
by: Feng, Guangnan, et al.
Published: (2024)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Wilkins: HPC In Situ Workflows Made Easy
by: Yildiz, Orcun, et al.
Published: (2024)
by: Yildiz, Orcun, et al.
Published: (2024)
Characterizing the Impact of Congestion in Modern HPC Interconnects
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
Similar Items
-
An Incremental Multi-Level, Multi-Scale Approach to Assessment of Multifidelity HPC Systems
by: Shilpika, Shilpika, et al.
Published: (2025) -
Towards Energy Efficient Co-Scheduling in HPC
by: Zheng, Zhong, et al.
Published: (2026) -
MRSch: Multi-Resource Scheduling for HPC
by: Li, Boyang, et al.
Published: (2024) -
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
by: Cornelius, Melanie, et al.
Published: (2025) -
Exploring Uncore Frequency Scaling for Heterogeneous Computing
by: Zheng, Zhong, et al.
Published: (2025)