Leveraging Hardware Performance Counters for Predicting Workload Interference in Vector Supercomputers
Fuente:
arXiv
Saved in:
| Main Authors: | Shubham, Takahashi, Keichi, Takizawa, Hiroyuki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
by: Takahashi, Keichi, et al.
Published: (2023)
by: Takahashi, Keichi, et al.
Published: (2023)
Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
by: Takahashi, Keichi, et al.
Published: (2024)
by: Takahashi, Keichi, et al.
Published: (2024)
Performance analysis of mdx II: A next-generation cloud platform for cross-disciplinary data science research
by: Takahashi, Keichi, et al.
Published: (2025)
by: Takahashi, Keichi, et al.
Published: (2025)
Otus Supercomputer
by: Ehtesabi, Sadaf, et al.
Published: (2025)
by: Ehtesabi, Sadaf, et al.
Published: (2025)
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)
by: Adefemi, Temitayo
Published: (2024)
TX-Digital Twin: Visualizing Supercomputer GPU Performance Data Stream
by: Baskakova, Elena, et al.
Published: (2026)
by: Baskakova, Elena, et al.
Published: (2026)
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads
by: Hu, Cunchen, et al.
Published: (2024)
by: Hu, Cunchen, et al.
Published: (2024)
Enabling Message Passing Interface Containers on the LUMI Supercomputer
by: Lazzaro, Alfio
Published: (2024)
by: Lazzaro, Alfio
Published: (2024)
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
by: Lu, Yao, et al.
Published: (2026)
by: Lu, Yao, et al.
Published: (2026)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
by: Ma, Xiaolong, et al.
Published: (2024)
by: Ma, Xiaolong, et al.
Published: (2024)
Scaling All-to-all Operations Across Emerging Many-Core Supercomputers
by: Kinkead, Shannon, et al.
Published: (2026)
by: Kinkead, Shannon, et al.
Published: (2026)
Tally: Non-Intrusive Performance Isolation for Concurrent Deep Learning Workloads
by: Zhao, Wei, et al.
Published: (2024)
by: Zhao, Wei, et al.
Published: (2024)
Supercomputer 3D Digital Twin for User Focused Real-Time Monitoring
by: Bergeron, William, et al.
Published: (2024)
by: Bergeron, William, et al.
Published: (2024)
Dispatching Odyssey: Exploring Performance in Computing Clusters under Real-world Workloads
by: Yildiz, Mert, et al.
Published: (2025)
by: Yildiz, Mert, et al.
Published: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
by: Esposito, Giuseppe, et al.
Published: (2025)
by: Esposito, Giuseppe, et al.
Published: (2025)
AAPA: An Archetype-Aware Predictive Autoscaler with Uncertainty Quantification for Serverless Workloads on Kubernetes
by: Zhang, Guilin, et al.
Published: (2025)
by: Zhang, Guilin, et al.
Published: (2025)
Supercomputing for High-speed Avoidance and Reactive Planning in Robots
by: Lachmansingh, Kieran S., et al.
Published: (2025)
by: Lachmansingh, Kieran S., et al.
Published: (2025)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
Study of Workload Interference with Intelligent Routing on Dragonfly
by: Kang, Yao, et al.
Published: (2024)
by: Kang, Yao, et al.
Published: (2024)
Agentic AI Workload Characteristics
by: Yuan, Yichao, et al.
Published: (2026)
by: Yuan, Yichao, et al.
Published: (2026)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
by: Zhang, Qiao, et al.
Published: (2025)
by: Zhang, Qiao, et al.
Published: (2025)
PWDFT-SW: Extending the Limit of Plane-Wave DFT Calculations to 16K Atoms on the New Sunway Supercomputer
by: Jiang, Qingcai, et al.
Published: (2024)
by: Jiang, Qingcai, et al.
Published: (2024)
TempoScale: A Cloud Workloads Prediction Approach Integrating Short-Term and Long-Term Information
by: Wen, Linfeng, et al.
Published: (2024)
by: Wen, Linfeng, et al.
Published: (2024)
Ksurf: Attention Kalman Filter and Principal Component Analysis for Prediction under Highly Variable Cloud Workloads
by: Dang'ana, Michael, et al.
Published: (2024)
by: Dang'ana, Michael, et al.
Published: (2024)
SPARS: A Reinforcement Learning-Enabled Simulator for Power Management in HPC Job Scheduling
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
by: Amrizal, Muhammad Alfian, et al.
Published: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
by: Go, Seokjin, et al.
Published: (2026)
by: Go, Seokjin, et al.
Published: (2026)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
by: Cornelius, Melanie, et al.
Published: (2025)
by: Cornelius, Melanie, et al.
Published: (2025)
AI Surrogate Model for Distributed Computing Workloads
by: Park, David K., et al.
Published: (2024)
by: Park, David K., et al.
Published: (2024)
Optimal Workload Placement on Multi-Instance GPUs
by: Turkkan, Bekir, et al.
Published: (2024)
by: Turkkan, Bekir, et al.
Published: (2024)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
by: Stoyanov, Radostin, et al.
Published: (2025)
by: Stoyanov, Radostin, et al.
Published: (2025)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
Scaling Large-scale GNN Training to Thousands of Processors on CPU-based Supercomputers
by: Zhuang, Chen, et al.
Published: (2024)
by: Zhuang, Chen, et al.
Published: (2024)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
Union: An Automatic Workload Manager for Accelerating Network Simulation
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
Distributed Load Balancing with Workload-Dependent Service Rates
by: Zhang, Wenxin, et al.
Published: (2024)
by: Zhang, Wenxin, et al.
Published: (2024)
Kub: Enabling Elastic HPC Workloads on Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
SYMPHONY: Improving Memory Management for LLM Inference Workloads
by: Agarwal, Saurabh, et al.
Published: (2024)
by: Agarwal, Saurabh, et al.
Published: (2024)
Workload Intelligence: Punching Holes Through the Cloud Abstraction
by: Huang, Lexiang, et al.
Published: (2024)
by: Huang, Lexiang, et al.
Published: (2024)
Similar Items
-
Performance Evaluation of a Next-Generation SX-Aurora TSUBASA Vector Supercomputer
by: Takahashi, Keichi, et al.
Published: (2023) -
Modernizing an Operational Real-time Tsunami Simulator to Support Diverse Hardware Platforms
by: Takahashi, Keichi, et al.
Published: (2024) -
Performance analysis of mdx II: A next-generation cloud platform for cross-disciplinary data science research
by: Takahashi, Keichi, et al.
Published: (2025) -
Otus Supercomputer
by: Ehtesabi, Sadaf, et al.
Published: (2025) -
Analysis of the Performance of the Matrix Multiplication Algorithm on the Cirrus Supercomputer
by: Adefemi, Temitayo
Published: (2024)