Characterizing Production GPU Workloads using System-wide Telemetry Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Cankur, Onur, Austin, Brian, Kulkarni, Dhruva, Bhatele, Abhinav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Pipit: Scripting the analysis of parallel execution traces
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023)
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024)
di: Cankur, Onur, et al.
Pubblicazione: (2024)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
di: Costello, Ian J., et al.
Pubblicazione: (2020)
di: Costello, Ian J., et al.
Pubblicazione: (2020)
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
di: Darzi, Erfan, et al.
Pubblicazione: (2025)
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
di: Davis, Joshua H., et al.
Pubblicazione: (2025)
di: Davis, Joshua H., et al.
Pubblicazione: (2025)
GPU Under Pressure: Estimating Application's Stress via Telemetry and Performance Counters
di: Esposito, Giuseppe, et al.
Pubblicazione: (2025)
di: Esposito, Giuseppe, et al.
Pubblicazione: (2025)
Taking GPU Programming Models to Task for Performance Portability
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
di: Davis, Joshua H., et al.
Pubblicazione: (2024)
CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads
di: Stoyanov, Radostin, et al.
Pubblicazione: (2025)
di: Stoyanov, Radostin, et al.
Pubblicazione: (2025)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
di: Nichols, Daniel, et al.
Pubblicazione: (2025)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
di: Davis, Joshua H., et al.
Pubblicazione: (2026)
ML-based Modeling to Predict I/O Performance on Different Storage Sub-systems
di: Xu, Yiheng, et al.
Pubblicazione: (2023)
di: Xu, Yiheng, et al.
Pubblicazione: (2023)
A Practical Two-Stage Framework for GPU Resource and Power Prediction in Heterogeneous HPC Systems
di: Oztop, Beste, et al.
Pubblicazione: (2026)
di: Oztop, Beste, et al.
Pubblicazione: (2026)
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
di: Xiang, Yuxing, et al.
Pubblicazione: (2025)
di: Xiang, Yuxing, et al.
Pubblicazione: (2025)
MERBIT: A GPU-Based SpMV Method for Iterative Workloads
di: Zhang, Qi, et al.
Pubblicazione: (2026)
di: Zhang, Qi, et al.
Pubblicazione: (2026)
Data Management System Analysis for Distributed Computing Workloads
di: Hsu, Kuan-Chieh, et al.
Pubblicazione: (2025)
di: Hsu, Kuan-Chieh, et al.
Pubblicazione: (2025)
PAL: A Variability-Aware Policy for Scheduling ML Workloads in GPU Clusters
di: Jain, Rutwik, et al.
Pubblicazione: (2024)
di: Jain, Rutwik, et al.
Pubblicazione: (2024)
PRISM: Dynamic Primitive-Based Forecasting for Large-Scale GPU Cluster Workloads
di: Wu, Xin, et al.
Pubblicazione: (2026)
di: Wu, Xin, et al.
Pubblicazione: (2026)
HPC-Coder: Modeling Parallel Programs using Large Language Models
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
di: Nichols, Daniel, et al.
Pubblicazione: (2023)
An Online Fragmentation-Aware Scheduler for Managing GPU-Sharing Workloads on Multi-Instance GPUs
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
di: Ting, Hsu-Tzu, et al.
Pubblicazione: (2025)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
di: Scheinert, Dominik, et al.
Pubblicazione: (2026)
di: Scheinert, Dominik, et al.
Pubblicazione: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
di: Parayil, Anjaly, et al.
Pubblicazione: (2024)
Scheduling Data-Intensive Workloads in Large-Scale Distributed Systems: Trends and Challenges
di: Stavrinides, Georgios L., et al.
Pubblicazione: (2025)
di: Stavrinides, Georgios L., et al.
Pubblicazione: (2025)
From Data Center IoT Telemetry to Data Analytics Chatbots -- Virtual Knowledge Graph is All You Need
di: Khan, Junaid Ahmed, et al.
Pubblicazione: (2025)
di: Khan, Junaid Ahmed, et al.
Pubblicazione: (2025)
Maya: Optimizing Deep Learning Training Workloads using GPU Runtime Emulation
di: Yarlagadda, Srihas, et al.
Pubblicazione: (2025)
di: Yarlagadda, Srihas, et al.
Pubblicazione: (2025)
A Framework for Fine-Grained Synchronization of Dependent GPU Kernels
di: Jangda, Abhinav, et al.
Pubblicazione: (2023)
di: Jangda, Abhinav, et al.
Pubblicazione: (2023)
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
di: Jain, Rutwik, et al.
Pubblicazione: (2026)
Crossword: Adaptive Consensus for Dynamic Data-Heavy Workloads
di: Hu, Guanzhou, et al.
Pubblicazione: (2025)
di: Hu, Guanzhou, et al.
Pubblicazione: (2025)
Eventually-Consistent Federated Scheduling for Data Center Workloads
di: Thiyyakat, Meghana, et al.
Pubblicazione: (2023)
di: Thiyyakat, Meghana, et al.
Pubblicazione: (2023)
MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
di: Xue, Chunyu, et al.
Pubblicazione: (2026)
Characterization-Guided GPU Fault Resilience in NVIDIA MPS
di: Liu, Rixin, et al.
Pubblicazione: (2026)
di: Liu, Rixin, et al.
Pubblicazione: (2026)
Understanding Data Movement in AMD Multi-GPU Systems with Infinity Fabric
di: Schieffer, Gabin, et al.
Pubblicazione: (2024)
di: Schieffer, Gabin, et al.
Pubblicazione: (2024)
Can Large Language Models Write Parallel Code?
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
di: Nichols, Daniel, et al.
Pubblicazione: (2024)
Ensemble Method for System Failure Detection Using Large-Scale Telemetry Data
di: Mudgal, Priyanka, et al.
Pubblicazione: (2024)
di: Mudgal, Priyanka, et al.
Pubblicazione: (2024)
Orchestrating Mixed-Criticality Cloud Workloads in Reconfigurable Manufacturing Systems
di: Barletta, Marco, et al.
Pubblicazione: (2024)
di: Barletta, Marco, et al.
Pubblicazione: (2024)
Profiling and Modeling of Power Characteristics of Leadership-Scale HPC System Workloads
di: Karimi, Ahmad Maroof, et al.
Pubblicazione: (2024)
di: Karimi, Ahmad Maroof, et al.
Pubblicazione: (2024)
Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model Training
di: Chen, Chang, et al.
Pubblicazione: (2025)
di: Chen, Chang, et al.
Pubblicazione: (2025)
Agentic AI Workload Characteristics
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
Pandemics In Silico: Scaling an Agent-Based Simulation on Realistic Social Contact Networks
di: Kitson, Joy, et al.
Pubblicazione: (2024)
di: Kitson, Joy, et al.
Pubblicazione: (2024)
Evaluating Malleable Job Scheduling in HPC Clusters using Real-World Workloads
di: Zojer, Patrick, et al.
Pubblicazione: (2026)
di: Zojer, Patrick, et al.
Pubblicazione: (2026)
SEE++: Evolving Snowpark Execution Environment for Modern Workloads
di: Jain, Gaurav, et al.
Pubblicazione: (2025)
di: Jain, Gaurav, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Pipit: Scripting the analysis of parallel execution traces
di: Bhatele, Abhinav, et al.
Pubblicazione: (2023) -
Automated Programmatic Performance Analysis of Parallel Programs
di: Cankur, Onur, et al.
Pubblicazione: (2024) -
Analytics of Longitudinal System Monitoring Data for Performance Prediction
di: Costello, Ian J., et al.
Pubblicazione: (2020) -
Host-Side Telemetry for Performance Diagnosis in Cloud and HPC GPU Infrastructure
di: Darzi, Erfan, et al.
Pubblicazione: (2025) -
ParEval-Repo: A Benchmark Suite for Evaluating LLMs with Repository-level HPC Translation Tasks
di: Davis, Joshua H., et al.
Pubblicazione: (2025)