Designing Datacenter Power Delivery Hierarchies for the AI Era
Fuente:
arXiv
Salvato in:
| Autori principali: | Wilkins, Grant, Kazhamiaka, Fiodar, Kumbhare, Alok Gautam, Zhang, Chaojie, Bianchini, Ricardo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
di: Wilkins, Grant, et al.
Pubblicazione: (2026)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025)
di: Choukse, Esha, et al.
Pubblicazione: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
di: Lettich, Francesco, et al.
Pubblicazione: (2024)
di: Lettich, Francesco, et al.
Pubblicazione: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
di: Oviedo, Felipe, et al.
Pubblicazione: (2025)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
di: Qiu, Haoran, et al.
Pubblicazione: (2026)
Hybrid Heterogeneous Clusters Can Lower the Energy Consumption of LLM Inference Workloads
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
Towards Resource-Efficient Compound AI Systems
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
di: Chaudhry, Gohar Irfan, et al.
Pubblicazione: (2025)
Intent-based System Design and Operation
di: Anand, Vaastav, et al.
Pubblicazione: (2025)
di: Anand, Vaastav, et al.
Pubblicazione: (2025)
The AI_INFN Platform: Artificial Intelligence Development in the Cloud
di: Anderlini, Lucio, et al.
Pubblicazione: (2025)
di: Anderlini, Lucio, et al.
Pubblicazione: (2025)
The (R)evolution of Scientific Workflows in the Agentic AI Era: Towards Autonomous Science
di: Shin, Woong, et al.
Pubblicazione: (2025)
di: Shin, Woong, et al.
Pubblicazione: (2025)
Datacenter Energy Optimized Power Profiles
di: Narayanaswamy, Sreedhar, et al.
Pubblicazione: (2025)
di: Narayanaswamy, Sreedhar, et al.
Pubblicazione: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
di: Stojkovic, Jovan, et al.
Pubblicazione: (2024)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
di: Mehboob, Talha, et al.
Pubblicazione: (2025)
di: Mehboob, Talha, et al.
Pubblicazione: (2025)
Sustainable Carbon-Aware and Water-Efficient LLM Scheduling in Geo-Distributed Cloud Datacenters
di: Moore, Hayden, et al.
Pubblicazione: (2025)
di: Moore, Hayden, et al.
Pubblicazione: (2025)
Fire-Flyer AI-HPC: A Cost-Effective Software-Hardware Co-Design for Deep Learning
di: An, Wei, et al.
Pubblicazione: (2024)
di: An, Wei, et al.
Pubblicazione: (2024)
Accelerating Latency-Critical Applications with AI-Powered Semi-Automatic Fine-Grained Parallelization on SMT Processors
di: Los, Denis, et al.
Pubblicazione: (2025)
di: Los, Denis, et al.
Pubblicazione: (2025)
StreamServe: Adaptive Speculative Flows for Low-Latency Disaggregated LLM Serving
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
di: Kumar, Satyam, et al.
Pubblicazione: (2026)
Scalable AI-assisted Workflow Management for Detector Design Optimization Using Distributed Computing
di: Anderson, Derek, et al.
Pubblicazione: (2026)
di: Anderson, Derek, et al.
Pubblicazione: (2026)
Autonomous Systems Dependability in the era of AI: Design Challenges in Safety, Security, Reliability and Certification
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
di: Ranjbar, Behnaz, et al.
Pubblicazione: (2026)
HybridEP: Scaling Expert Parallelism to Cross-Datacenter Scenario via Hybrid Expert/Data Transmission
di: Yang, Weihao, et al.
Pubblicazione: (2025)
di: Yang, Weihao, et al.
Pubblicazione: (2025)
DCGen 1.1 Technical Report: Generating Datacenter Configurations (including IT, Power, Cooling)
di: Gnibga, Wedan Emmanuel, et al.
Pubblicazione: (2026)
di: Gnibga, Wedan Emmanuel, et al.
Pubblicazione: (2026)
From Barrier to Bridge: The Case for AI Data Center/Power Grid Co-Design
di: Bashir, Noman, et al.
Pubblicazione: (2026)
di: Bashir, Noman, et al.
Pubblicazione: (2026)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
di: Wilkins, Grant, et al.
Pubblicazione: (2024)
When AI Bends Metal: AI-Assisted Optimization of Design Parameters in Sheet Metal Forming
di: Tarraf, Ahmad, et al.
Pubblicazione: (2025)
di: Tarraf, Ahmad, et al.
Pubblicazione: (2025)
PALS: Power-Aware LLM Serving for Mixture-of-Experts Models
di: Hankendi, Can, et al.
Pubblicazione: (2026)
di: Hankendi, Can, et al.
Pubblicazione: (2026)
Deploying Foundation Model Powered Agent Services: A Survey
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
di: Xu, Wenchao, et al.
Pubblicazione: (2024)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
di: Yuan, Yichao, et al.
Pubblicazione: (2026)
HFX: Joint Design of Algorithms and Systems for Multi-SLO Serving and Fast Scaling
di: Yousefijamarani, Zahra, et al.
Pubblicazione: (2025)
di: Yousefijamarani, Zahra, et al.
Pubblicazione: (2025)
Xe-Forge: Multi-Stage LLM-Powered Kernel Optimization for Intel GPU
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
di: Spoczynski, Marcin, et al.
Pubblicazione: (2026)
Scaling Performance of Large Language Model Pretraining
di: Interrante-Grant, Alexander, et al.
Pubblicazione: (2025)
di: Interrante-Grant, Alexander, et al.
Pubblicazione: (2025)
AI Benchmarks and Datasets for LLM Evaluation
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
di: Ivanov, Todor, et al.
Pubblicazione: (2024)
The Case for Co-Designing Model Architectures with Hardware
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
di: Anthony, Quentin, et al.
Pubblicazione: (2024)
Serving Compound Inference Systems on Datacenter GPUs
di: Devata, Sriram, et al.
Pubblicazione: (2026)
di: Devata, Sriram, et al.
Pubblicazione: (2026)
The infrastructure powering IBM's Gen AI model development
di: Gershon, Talia, et al.
Pubblicazione: (2024)
di: Gershon, Talia, et al.
Pubblicazione: (2024)
Adaptation of AI-accelerated CFD Simulations to the IPU platform
di: Rosciszewski, P., et al.
Pubblicazione: (2026)
di: Rosciszewski, P., et al.
Pubblicazione: (2026)
AI Factories: It's time to rethink the Cloud-HPC divide
di: Lopez, Pedro Garcia, et al.
Pubblicazione: (2025)
di: Lopez, Pedro Garcia, et al.
Pubblicazione: (2025)
MSCCL++: Rethinking GPU Communication Abstractions for AI Inference
di: Hwang, Changho, et al.
Pubblicazione: (2025)
di: Hwang, Changho, et al.
Pubblicazione: (2025)
Decentralized AI: Permissionless LLM Inference on POKT Network
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
di: Olshansky, Daniel, et al.
Pubblicazione: (2024)
Documenti analoghi
-
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
di: Wilkins, Grant, et al.
Pubblicazione: (2026) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
Power Stabilization for AI Training Datacenters
di: Choukse, Esha, et al.
Pubblicazione: (2025) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
di: Stojkovic, Jovan, et al.
Pubblicazione: (2025) -
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
di: Lettich, Francesco, et al.
Pubblicazione: (2024)