The Energy Cost of Execution-Idle in GPU Clusters
Fuente:
arXiv
Guardado en:
| Autores principales: | Lei, Yiran, Fernandez, Jared, Kypriotis, Vasilis, Skarlatos, Dimitrios, Strubell, Emma, Sherry, Justine, Vosler, Daniel |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026)
por: Jain, Rutwik, et al.
Publicado: (2026)
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024)
por: Wang, Yidi, et al.
Publicado: (2024)
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026)
por: Davis, Joshua H., et al.
Publicado: (2026)
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025)
por: Lahiry, Ankur, et al.
Publicado: (2025)
On the Partitioning of GPU Power among Multi-Instances
por: Vamja, Tirth, et al.
Publicado: (2025)
por: Vamja, Tirth, et al.
Publicado: (2025)
Towards Portability at Scale: A Cross-Architecture Performance Evaluation of a GPU-enabled Shallow Water Solver
por: Villalobos, Johansell, et al.
Publicado: (2025)
por: Villalobos, Johansell, et al.
Publicado: (2025)
Green or Fast? Learning to Balance Cold Starts and Idle Carbon in Serverless Computing
por: Sun, Bowen, et al.
Publicado: (2026)
por: Sun, Bowen, et al.
Publicado: (2026)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
por: Afzal, Ayesha, et al.
Publicado: (2024)
por: Afzal, Ayesha, et al.
Publicado: (2024)
Disaggregated Design for GPU-Based Volumetric Data Structures
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
por: Meneghin, Massimiliano, et al.
Publicado: (2025)
Taking GPU Programming Models to Task for Performance Portability
por: Davis, Joshua H., et al.
Publicado: (2024)
por: Davis, Joshua H., et al.
Publicado: (2024)
Accelerating Mobile Inference through Fine-Grained CPU-GPU Co-Execution
por: Li, Zhuojin, et al.
Publicado: (2025)
por: Li, Zhuojin, et al.
Publicado: (2025)
Profiling and optimization of multi-card GPU machine learning jobs
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
CUTHERMO: Understanding GPU Memory Inefficiencies with Heat Map Profiling
por: Zhao, Yanbo, et al.
Publicado: (2025)
por: Zhao, Yanbo, et al.
Publicado: (2025)
FACT: Compositional Kernel Synthesis with a Three-Stage Agentic Workflow
por: Heidari, Sina, et al.
Publicado: (2026)
por: Heidari, Sina, et al.
Publicado: (2026)
High-Performance Portable GPU Primitives for Arbitrary Types and Operators in Julia
por: Pilliat, Emmanuel
Publicado: (2026)
por: Pilliat, Emmanuel
Publicado: (2026)
Efficient allocation of image recognition and LLM tasks on multi-GPU system
por: Lawenda, Marcin, et al.
Publicado: (2025)
por: Lawenda, Marcin, et al.
Publicado: (2025)
Data-Driven Analysis to Understand GPU Hardware Resource Usage of Optimizations
por: Islam, Tanzima Z., et al.
Publicado: (2024)
por: Islam, Tanzima Z., et al.
Publicado: (2024)
GPU Cluster Scheduling for Network-Sensitive Deep Learning
por: Sharma, Aakash, et al.
Publicado: (2024)
por: Sharma, Aakash, et al.
Publicado: (2024)
Efficient GPU-Centered Singular Value Decomposition Using the Divide-and-Conquer Method
por: Liu, Shifang, et al.
Publicado: (2025)
por: Liu, Shifang, et al.
Publicado: (2025)
Less is More: Optimizing Function Calling for LLM Execution on Edge Devices
por: Paramanayakam, Varatheepan, et al.
Publicado: (2024)
por: Paramanayakam, Varatheepan, et al.
Publicado: (2024)
Dissecting CPU-GPU Unified Physical Memory on AMD MI300A APUs
por: Wahlgren, Jacob, et al.
Publicado: (2025)
por: Wahlgren, Jacob, et al.
Publicado: (2025)
LEO: Tracing GPU Stall Root Causes via Cross-Vendor Backward Slicing
por: Xia, Yuning, et al.
Publicado: (2026)
por: Xia, Yuning, et al.
Publicado: (2026)
Fast and Scalable Mixed Precision Euclidean Distance Calculations Using GPU Tensor Cores
por: Curless, Brian, et al.
Publicado: (2025)
por: Curless, Brian, et al.
Publicado: (2025)
HybridGen: Efficient LLM Generative Inference via CPU-GPU Hybrid Computing
por: Lin, Mao, et al.
Publicado: (2026)
por: Lin, Mao, et al.
Publicado: (2026)
xMem: A CPU-Based Approach for Accurate Estimation of GPU Memory in Deep Learning Training Workloads
por: Shi, Jiabo, et al.
Publicado: (2025)
por: Shi, Jiabo, et al.
Publicado: (2025)
Ecomap: Sustainability-Driven Optimization of Multi-Tenant DNN Execution on Edge Servers
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
por: Paramanayakam, Varatheepan, et al.
Publicado: (2025)
Node Compass: Multilevel Tracing and Debugging of Request Executions in JavaScript-Based Web-Servers
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
por: Kabamba, Herve Mbikayi, et al.
Publicado: (2023)
Evaluating HPC-Style CPU Performance and Cost in Virtualized Cloud Infrastructures
por: Tharwani, Jay, et al.
Publicado: (2025)
por: Tharwani, Jay, et al.
Publicado: (2025)
Cost-Performance Evaluation of General Compute Instances: AWS, Azure, GCP, and OCI
por: Tharwani, Jay, et al.
Publicado: (2024)
por: Tharwani, Jay, et al.
Publicado: (2024)
Energy-Aware Computing in the Year 2026
por: Tchakoute, Roblex Nana, et al.
Publicado: (2026)
por: Tchakoute, Roblex Nana, et al.
Publicado: (2026)
"Two-Stagification": Job Dispatching in Large-Scale Clusters via a Two-Stage Architecture
por: Yildiz, Mert, et al.
Publicado: (2025)
por: Yildiz, Mert, et al.
Publicado: (2025)
Extracting Practical, Actionable Energy Insights from Supercomputer Telemetry and Logs
por: Cornelius, Melanie, et al.
Publicado: (2025)
por: Cornelius, Melanie, et al.
Publicado: (2025)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
por: Dutt, Anurag, et al.
Publicado: (2025)
por: Dutt, Anurag, et al.
Publicado: (2025)
A Comprehensive Analysis of Process Energy Consumption on Multi-Socket Systems with GPUs
por: León-Vega, Luis G., et al.
Publicado: (2024)
por: León-Vega, Luis G., et al.
Publicado: (2024)
Integrating Performance Tools in Model Reasoning for GPU Kernel Optimization
por: Nichols, Daniel, et al.
Publicado: (2025)
por: Nichols, Daniel, et al.
Publicado: (2025)
Arkade: k-Nearest Neighbor Search With Non-Euclidean Distances using GPU Ray Tracing
por: Mandarapu, Durga, et al.
Publicado: (2023)
por: Mandarapu, Durga, et al.
Publicado: (2023)
The Landscape of GPU-Centric Communication
por: Unat, Didem, et al.
Publicado: (2024)
por: Unat, Didem, et al.
Publicado: (2024)
GigaAPI for GPU Parallelization
por: Suvarna, M., et al.
Publicado: (2025)
por: Suvarna, M., et al.
Publicado: (2025)
Parallelizing a modern GPU simulator
por: Huerta, Rodrigo, et al.
Publicado: (2025)
por: Huerta, Rodrigo, et al.
Publicado: (2025)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
por: Maczan, Jędrzej
Publicado: (2026)
por: Maczan, Jędrzej
Publicado: (2026)
Ejemplares similares
-
Minos: Systematically Classifying Performance and Power Characteristics of GPU Workloads on HPC Clusters
por: Jain, Rutwik, et al.
Publicado: (2026) -
Unleashing the Power of Preemptive Priority-based Scheduling for Real-Time GPU Tasks
por: Wang, Yidi, et al.
Publicado: (2024) -
KEET: Explaining Performance of GPU Kernels Using LLM Agents
por: Davis, Joshua H., et al.
Publicado: (2026) -
Scalable GPU Performance Variability Analysis framework
por: Lahiry, Ankur, et al.
Publicado: (2025) -
On the Partitioning of GPU Power among Multi-Instances
por: Vamja, Tirth, et al.
Publicado: (2025)