Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
Fuente:
arXiv
Saved in:
| Main Authors: | Darzi, Erfan, Pareja, Aldo, Bharadwaj, Shreeanant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Predictable LLM Serving on GPU Clusters
by: Darzi, Erfan, et al.
Published: (2025)
by: Darzi, Erfan, et al.
Published: (2025)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025)
by: Tharwani, Jay, et al.
Published: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
by: Esposito, Giuseppe, et al.
Published: (2025)
by: Esposito, Giuseppe, et al.
Published: (2025)
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025)
by: Cankur, Onur, et al.
Published: (2025)
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)
by: Chazapis, Antony, et al.
Published: (2024)
GICC: A High-Performance Runtime for GPU-Initiated Communication and Coordination in Modern HPC Systems
by: Shan, Baodi, et al.
Published: (2026)
by: Shan, Baodi, et al.
Published: (2026)
GPU-Accelerated Distributed QAOA on Large-scale HPC Ecosystems
by: Xu, Zhihao, et al.
Published: (2025)
by: Xu, Zhihao, et al.
Published: (2025)
IOAgent: Democratizing Trustworthy HPC I/O Performance Diagnosis Capability via LLMs
by: Egersdoerfer, Chris, et al.
Published: (2026)
by: Egersdoerfer, Chris, et al.
Published: (2026)
An Elastic Job Scheduler for HPC Applications on the Cloud
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Sarus Suite: Cloud-native Containers for HPC
by: Madonna, Alberto, et al.
Published: (2026)
by: Madonna, Alberto, et al.
Published: (2026)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
by: Jain, Rutwik, et al.
Published: (2026)
by: Jain, Rutwik, et al.
Published: (2026)
Towards an Adaptive Runtime System for Cloud-Native HPC
by: Bhosale, Aditya, et al.
Published: (2026)
by: Bhosale, Aditya, et al.
Published: (2026)
HPC resources for CMS offline computing: An integration and scalability challenge for the Submission Infrastructure
by: Yzquierdo, Antonio Perez-Calero, et al.
Published: (2024)
by: Yzquierdo, Antonio Perez-Calero, et al.
Published: (2024)
Incisor: Ex Ante Cloud Instance Selection for HPC Jobs
by: Laurenzano, Michael A., et al.
Published: (2026)
by: Laurenzano, Michael A., et al.
Published: (2026)
Understanding Layered Portability from HPC to Cloud in Containerized Environments
by: Medeiros, Daniel, et al.
Published: (2024)
by: Medeiros, Daniel, et al.
Published: (2024)
BSODiag: A Global Diagnosis Framework for Batch Servers Outage in Large-scale Cloud Infrastructure Systems
by: Duan, Tao, et al.
Published: (2025)
by: Duan, Tao, et al.
Published: (2025)
Usability Evaluation of Cloud for HPC Applications
by: Sochat, Vanessa, et al.
Published: (2025)
by: Sochat, Vanessa, et al.
Published: (2025)
Harnessing Integrated CPU-GPU System Memory for HPC: a first look into Grace Hopper
by: Schieffer, Gabin, et al.
Published: (2024)
by: Schieffer, Gabin, et al.
Published: (2024)
HPCAdvisor: A Tool for Assisting Users in Selecting HPC Resources in the Cloud
by: Netto, Marco A. S.
Published: (2024)
by: Netto, Marco A. S.
Published: (2024)
Federated Single Sign-On and Zero Trust Co-design for AI and HPC Digital Research Infrastructures
by: Alam, Sadaf R., et al.
Published: (2024)
by: Alam, Sadaf R., et al.
Published: (2024)
OpenFLAME: A Federated Spatial Naming Infrastructure
by: Bharadwaj, Sagar, et al.
Published: (2024)
by: Bharadwaj, Sagar, et al.
Published: (2024)
Characterizing the Impact of Congestion in Modern HPC Interconnects
by: Piarulli, Lorenzo, et al.
Published: (2026)
by: Piarulli, Lorenzo, et al.
Published: (2026)
AI-coupled HPC Workflow Applications, Middleware and Performance
by: Brewer, Wes, et al.
Published: (2024)
by: Brewer, Wes, et al.
Published: (2024)
Performance comparison of Dask and Apache Spark on HPC systems for Neuroimaging
by: Dugré, Mathieu, et al.
Published: (2024)
by: Dugré, Mathieu, et al.
Published: (2024)
Computational Performance and Energy Efficiency of ARM based HPC servers
by: Schirmer, Oskar
Published: (2024)
by: Schirmer, Oskar
Published: (2024)
ParvaGPU: Efficient Spatial GPU Sharing for Large-Scale DNN Inference in Cloud Environments
by: Lee, Munkyu, et al.
Published: (2024)
by: Lee, Munkyu, et al.
Published: (2024)
Agora: Bridging the GPU Cloud Resource-Price Disconnect
by: McDougall, Ian, et al.
Published: (2025)
by: McDougall, Ian, et al.
Published: (2025)
DECICE: AI-Driven Scheduling and Digital Twin Integration for the Cloud-HPC-Edge Compute Continuum
by: Sharma, Aasish Kumar, et al.
Published: (2026)
by: Sharma, Aasish Kumar, et al.
Published: (2026)
Multiple Sides of 36 Coins: Measuring Peer-to-Peer Infrastructure Across Cryptocurrencies
by: Kiffer, Lucianna, et al.
Published: (2025)
by: Kiffer, Lucianna, et al.
Published: (2025)
Optimization Opportunities for Cloud-Based Data Pipeline Infrastructures
by: Jablonski, Johannes, et al.
Published: (2026)
by: Jablonski, Johannes, et al.
Published: (2026)
An Analysis of HPC and Edge Architectures in the Cloud
by: Santillan, Steven, et al.
Published: (2025)
by: Santillan, Steven, et al.
Published: (2025)
A Performance Analysis of Task Scheduling for UQ Workflows on HPC Systems
by: Loi, Chung Ming, et al.
Published: (2025)
by: Loi, Chung Ming, et al.
Published: (2025)
Integrating and Characterizing HPC Task Runtime Systems for hybrid AI-HPC workloads
by: Merzky, Andre, et al.
Published: (2025)
by: Merzky, Andre, et al.
Published: (2025)
ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace
by: Shi, Ruimin, et al.
Published: (2025)
by: Shi, Ruimin, et al.
Published: (2025)
Performance characterisation of the 64-core SG2042 RISC-V CPU for HPC
by: Brown, Nick, et al.
Published: (2024)
by: Brown, Nick, et al.
Published: (2024)
SuperBench: Improving Cloud AI Infrastructure Reliability with Proactive Validation
by: Xiong, Yifan, et al.
Published: (2024)
by: Xiong, Yifan, et al.
Published: (2024)
A Multi-Objective Framework for Optimizing GPU-Enabled VM Placement in Cloud Data Centers with Multi-Instance GPU Technology
by: Siavashi, Ahmad, et al.
Published: (2025)
by: Siavashi, Ahmad, et al.
Published: (2025)
Application Experiences on a GPU-Accelerated Arm-based HPC Testbed
by: Elwasif, Wael, et al.
Published: (2022)
by: Elwasif, Wael, et al.
Published: (2022)
Analysis of the carbon footprint of HPC
by: Benhari, Abdessalam, et al.
Published: (2025)
by: Benhari, Abdessalam, et al.
Published: (2025)
HPC with Enhanced User Separation
by: Prout, Andrew, et al.
Published: (2024)
by: Prout, Andrew, et al.
Published: (2024)
Similar Items
-
Predictable LLM Serving on GPU Clusters
by: Darzi, Erfan, et al.
Published: (2025) -
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
by: Tharwani, Jay, et al.
Published: (2025) -
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
by: Esposito, Giuseppe, et al.
Published: (2025) -
Characterizing Production GPU Workloads using System-wide Telemetry Data
by: Cankur, Onur, et al.
Published: (2025) -
Running Cloud-native Workloads on HPC with High-Performance Kubernetes
by: Chazapis, Antony, et al.
Published: (2024)