Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
Fuente:
arXiv
Salvato in:
| Autori principali: | Chu, Xiaoyu, Hofstätter, Daniel, Ilager, Shashikant, Talluri, Sacheendra, Kampert, Duncan, Podareanu, Damian, Duplyakin, Dmitry, Brandic, Ivona, Iosup, Alexandru |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2026)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2026)
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
di: Maliakel, Paul Joe, et al.
Pubblicazione: (2025)
di: Maliakel, Paul Joe, et al.
Pubblicazione: (2025)
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
di: Nicolae, Radu, et al.
Pubblicazione: (2026)
di: Nicolae, Radu, et al.
Pubblicazione: (2026)
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
di: Chu, Xiaoyu, et al.
Pubblicazione: (2025)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2025)
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
di: Niewenhuis, Dante, et al.
Pubblicazione: (2026)
di: Niewenhuis, Dante, et al.
Pubblicazione: (2026)
ABBA-VSM: Time Series Classification using Symbolic Representation on the Edge
di: Kanatbekova, Meerzhan, et al.
Pubblicazione: (2024)
di: Kanatbekova, Meerzhan, et al.
Pubblicazione: (2024)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
di: Lei, Jianlong, et al.
Pubblicazione: (2026)
FLIGAN: Enhancing Federated Learning with Incomplete Data using GAN
di: Maliakel, Paul Joe, et al.
Pubblicazione: (2024)
di: Maliakel, Paul Joe, et al.
Pubblicazione: (2024)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
di: Ilager, Shashikant, et al.
Pubblicazione: (2025)
di: Ilager, Shashikant, et al.
Pubblicazione: (2025)
A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
di: Ilager, Shashikant, et al.
Pubblicazione: (2024)
di: Ilager, Shashikant, et al.
Pubblicazione: (2024)
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
di: May, Daniel, et al.
Pubblicazione: (2024)
di: May, Daniel, et al.
Pubblicazione: (2024)
Towards An Approach to Identify Divergences in Hardware Designs for HPC Workloads
di: Popovici, Doru Thom, et al.
Pubblicazione: (2025)
di: Popovici, Doru Thom, et al.
Pubblicazione: (2025)
EasyRider: Mitigating Power Transients in Datacenter-Scale Training Workloads
di: Jensen, Dillon, et al.
Pubblicazione: (2026)
di: Jensen, Dillon, et al.
Pubblicazione: (2026)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
di: Zilic, Josip, et al.
Pubblicazione: (2024)
di: Zilic, Josip, et al.
Pubblicazione: (2024)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
di: Karami, Rachid, et al.
Pubblicazione: (2024)
di: Karami, Rachid, et al.
Pubblicazione: (2024)
Cloud Uptime Archive: Open-Access Availability Data of Web, Cloud, and Gaming Services
di: Talluri, Sacheendra, et al.
Pubblicazione: (2025)
di: Talluri, Sacheendra, et al.
Pubblicazione: (2025)
DDC: A Vision for a Disaggregated Datacenter
di: Ewais, Mohammad, et al.
Pubblicazione: (2024)
di: Ewais, Mohammad, et al.
Pubblicazione: (2024)
Scalable and RISC-V Programmable Near-Memory Computing Architectures for Edge Nodes
di: Caon, Michele, et al.
Pubblicazione: (2024)
di: Caon, Michele, et al.
Pubblicazione: (2024)
Empirical Measurements of AI Training Power Demand on a GPU-Accelerated Node
di: Latif, Imran, et al.
Pubblicazione: (2024)
di: Latif, Imran, et al.
Pubblicazione: (2024)
A Node-Based Polar List Decoder with Frame Interleaving and Ensemble Decoding Support
di: Ren, Yuqing, et al.
Pubblicazione: (2024)
di: Ren, Yuqing, et al.
Pubblicazione: (2024)
Physical Design Exploration of a Wire-Friendly Domain-Specific Processor for Angstrom-Era Nodes
di: Ruotolo, Lorenzo, et al.
Pubblicazione: (2025)
di: Ruotolo, Lorenzo, et al.
Pubblicazione: (2025)
Empirically-Calibrated H100 Node Power Models for Reducing Uncertainty in AI Training Energy Estimation
di: Newkirk, Alex C., et al.
Pubblicazione: (2025)
di: Newkirk, Alex C., et al.
Pubblicazione: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
di: Ziller, Thomas, et al.
Pubblicazione: (2026)
On Reducing the Execution Latency of Superconducting Quantum Processors via Quantum Job Scheduling
di: Wu, Wenjie, et al.
Pubblicazione: (2024)
di: Wu, Wenjie, et al.
Pubblicazione: (2024)
Workload Characterization for Branch Predictability
di: Vikas, FNU, et al.
Pubblicazione: (2025)
di: Vikas, FNU, et al.
Pubblicazione: (2025)
AgilePkgC: An Agile System Idle State Architecture for Energy Proportional Datacenter Servers
di: Antoniou, Georgia, et al.
Pubblicazione: (2022)
di: Antoniou, Georgia, et al.
Pubblicazione: (2022)
A4: Microarchitecture-Aware LLC Management for Datacenter Servers with Emerging I/O Devices
di: Park, Haneul, et al.
Pubblicazione: (2025)
di: Park, Haneul, et al.
Pubblicazione: (2025)
How Much Progress Has There Been in NVIDIA Datacenter GPUs?
di: Del Sozzo, Emanuele, et al.
Pubblicazione: (2026)
di: Del Sozzo, Emanuele, et al.
Pubblicazione: (2026)
Cross-layer Modeling and Design of Content Addressable Memories in Advanced Technology Nodes for Similarity Search
di: Narla, Siri, et al.
Pubblicazione: (2024)
di: Narla, Siri, et al.
Pubblicazione: (2024)
Toward Open-Source Chiplets for HPC and AI: Occamy and Beyond
di: Scheffler, Paul, et al.
Pubblicazione: (2025)
di: Scheffler, Paul, et al.
Pubblicazione: (2025)
Educating for Hardware Specialization in the Chiplet Era: A Path for the HPC Community
di: Yoshii, Kazutomo, et al.
Pubblicazione: (2024)
di: Yoshii, Kazutomo, et al.
Pubblicazione: (2024)
Affordable HPC: Leveraging Small Clusters for Big Data and Graph Computing
di: Wu, Ruilong, et al.
Pubblicazione: (2024)
di: Wu, Ruilong, et al.
Pubblicazione: (2024)
Enabling Time-Aware Priority Traffic Management over Distributed FPGA Nodes
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
di: Scionti, Alberto, et al.
Pubblicazione: (2025)
Scaling Multi-Node Mixture-of-Experts Inference Using Expert Activation Patterns
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
di: Bambhaniya, Abhimanyu, et al.
Pubblicazione: (2026)
Robotics Under Construction: Challenges on Job Sites
di: Uchiito, Haruki, et al.
Pubblicazione: (2025)
di: Uchiito, Haruki, et al.
Pubblicazione: (2025)
Architectural Classification of XR Workloads: Cross-Layer Archetypes and Implications
di: Shi, Xinyu, et al.
Pubblicazione: (2026)
di: Shi, Xinyu, et al.
Pubblicazione: (2026)
Allspark: Workload Orchestration for Visual Transformers on Processing In-Memory Systems
di: Ge, Mengke, et al.
Pubblicazione: (2024)
di: Ge, Mengke, et al.
Pubblicazione: (2024)
SDT: Cutting Datacenter Tax Through Simultaneous Data-Delivery Threads
di: Mamandipoor, Amin, et al.
Pubblicazione: (2025)
di: Mamandipoor, Amin, et al.
Pubblicazione: (2025)
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025)
di: Choukse, Esha, et al.
Pubblicazione: (2025)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
di: McDaniel, Adam, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Leveraging LLMs for Structured Information Extraction and Analysis from Cloud Incident Reports (Work In Progress Paper)
di: Chu, Xiaoyu, et al.
Pubblicazione: (2026) -
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
di: Maliakel, Paul Joe, et al.
Pubblicazione: (2025) -
M3SA: Exploring Datacenter Performance and Climate-Impact with Multi- and Meta-Model Simulation and Analysis
di: Nicolae, Radu, et al.
Pubblicazione: (2026) -
An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models
di: Chu, Xiaoyu, et al.
Pubblicazione: (2025) -
OpenDC-STEAM: Realistic Modeling and Systematic Exploration of Composable Techniques for Sustainable Datacenters
di: Niewenhuis, Dante, et al.
Pubblicazione: (2026)