Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
Fuente:
arXiv
Saved in:
| Main Authors: | Oviedo, Felipe, Kazhamiaka, Fiodar, Choukse, Esha, Kim, Allen, Luers, Amy, Nakagawa, Melanie, Bianchini, Ricardo, Ferres, Juan M. Lavista |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
by: Wilkins, Grant, et al.
Published: (2026)
by: Wilkins, Grant, et al.
Published: (2026)
Designing Datacenter Power Delivery Hierarchies for the AI Era
by: Wilkins, Grant, et al.
Published: (2026)
by: Wilkins, Grant, et al.
Published: (2026)
EcoServe: Designing Carbon-Aware AI Inference Systems
by: Li, Yueying, et al.
Published: (2025)
by: Li, Yueying, et al.
Published: (2025)
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Towards Resource-Efficient Compound AI Systems
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
StreamWise: Serving Multi-Modal Generation in Real-Time at Scale
by: Qiu, Haoran, et al.
Published: (2026)
by: Qiu, Haoran, et al.
Published: (2026)
Junctiond: Extending FaaS Runtimes with Kernel-Bypass
by: Saurez, Enrique, et al.
Published: (2024)
by: Saurez, Enrique, et al.
Published: (2024)
No Request Left Behind: Tackling Heterogeneity in Long-Context LLM Inference with Medha
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Computational Performance and Energy Efficiency of ARM based HPC servers
by: Schirmer, Oskar
Published: (2024)
by: Schirmer, Oskar
Published: (2024)
Optimizing Resource Allocation and Energy Efficiency in Federated Fog Computing for IoT
by: Shah, Syed Sarmad, et al.
Published: (2025)
by: Shah, Syed Sarmad, et al.
Published: (2025)
Energy Efficiency Support for Software Defined Networks: a Serverless Computing Approach
by: Banaie, Fatemeh, et al.
Published: (2024)
by: Banaie, Fatemeh, et al.
Published: (2024)
Cloud abstractions for AI workloads
by: Canini, Marco, et al.
Published: (2025)
by: Canini, Marco, et al.
Published: (2025)
The 1/W Law: An Analytical Study of Context-Length Routing Topology and GPU Generation Gains for LLM Inference Energy Efficiency
by: Chen, Huamin, et al.
Published: (2026)
by: Chen, Huamin, et al.
Published: (2026)
KV Cache Compression for Inference Efficiency in LLMs: A Review
by: Liu, Yanyu, et al.
Published: (2025)
by: Liu, Yanyu, et al.
Published: (2025)
CALVO: Improve Serving Efficiency for LLM Inferences with Intense Network Demands
by: Wang, Weiye, et al.
Published: (2026)
by: Wang, Weiye, et al.
Published: (2026)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
by: Malekzadeh, Mohammad, et al.
Published: (2023)
by: Malekzadeh, Mohammad, et al.
Published: (2023)
GreenFaaS: Maximizing Energy Efficiency of HPC Workloads with FaaS
by: Kamatar, Alok, et al.
Published: (2024)
by: Kamatar, Alok, et al.
Published: (2024)
ModServe: Modality- and Stage-Aware Resource Disaggregation for Scalable Multimodal Model Serving
by: Qiu, Haoran, et al.
Published: (2025)
by: Qiu, Haoran, et al.
Published: (2025)
Distributed On-Device LLM Inference With Over-the-Air Computation
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
WANSpec: Leveraging Global Compute Capacity for LLM Inference
by: Martin, Noah, et al.
Published: (2026)
by: Martin, Noah, et al.
Published: (2026)
DAK: Direct-Access-Enabled GPU Memory Offloading with Optimal Efficiency for LLM Inference
by: Lin, Shouxu, et al.
Published: (2026)
by: Lin, Shouxu, et al.
Published: (2026)
OSGym: Scalable OS Infra for Computer Use Agents
by: Qin, Zengyi, et al.
Published: (2025)
by: Qin, Zengyi, et al.
Published: (2025)
Exploring the Efficiency of Renewable Energy-based Modular Data Centers at Scale
by: Sun, Jinghan, et al.
Published: (2024)
by: Sun, Jinghan, et al.
Published: (2024)
Enhancing Energy Efficiency in Scientific Workflows through CFD based PIVAEs
by: Zahir, Ali, et al.
Published: (2026)
by: Zahir, Ali, et al.
Published: (2026)
Exploring the Frontiers of Energy Efficiency using Power Management at System Scale
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Decentralized LLM Inference over Edge Networks with Energy Harvesting
by: Khoshsirat, Aria, et al.
Published: (2024)
by: Khoshsirat, Aria, et al.
Published: (2024)
Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
by: Moghadampanah, Mona, et al.
Published: (2025)
by: Moghadampanah, Mona, et al.
Published: (2025)
Driving Computational Efficiency in Large-Scale Platforms using HPC Technologies
by: Mendez, Alexander Martinez, et al.
Published: (2026)
by: Mendez, Alexander Martinez, et al.
Published: (2026)
Jiagu: Optimizing Serverless Computing Resource Utilization with Harmonized Efficiency and Practicability
by: Liu, Qingyuan, et al.
Published: (2024)
by: Liu, Qingyuan, et al.
Published: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Offline Energy-Optimal LLM Serving: Workload-Based Energy Models for LLM Inference on Heterogeneous Systems
by: Wilkins, Grant, et al.
Published: (2024)
by: Wilkins, Grant, et al.
Published: (2024)
FaasMeter: Energy-First Serverless Computing
by: Rehman, Abdul, et al.
Published: (2024)
by: Rehman, Abdul, et al.
Published: (2024)
Hybrid Cloud Architectures for Research Computing: Applications and Use Cases
by: Stiensmeier, Xaver, et al.
Published: (2026)
by: Stiensmeier, Xaver, et al.
Published: (2026)
Supervised Distributed Computing: Efficiency and Robustness under a Majority of Adversarial Workers
by: Augustine, John, et al.
Published: (2026)
by: Augustine, John, et al.
Published: (2026)
Comparative Analysis of Lightweight Kubernetes Distributions for Edge Computing: Performance and Resource Efficiency
by: Yakubov, Diyaz, et al.
Published: (2025)
by: Yakubov, Diyaz, et al.
Published: (2025)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
by: Özcan, Miray, et al.
Published: (2025)
by: Özcan, Miray, et al.
Published: (2025)
Intelligent Router for LLM Workloads: Improving Performance Through Workload-Aware Load Balancing
by: Jain, Kunal, et al.
Published: (2024)
by: Jain, Kunal, et al.
Published: (2024)
EdgeShard: Efficient LLM Inference via Collaborative Edge Computing
by: Zhang, Mingjin, et al.
Published: (2024)
by: Zhang, Mingjin, et al.
Published: (2024)
Similar Items
-
From Servers to Sites: Compositional Power Trace Generation of LLM Inference for Infrastructure Planning
by: Wilkins, Grant, et al.
Published: (2026) -
Designing Datacenter Power Delivery Hierarchies for the AI Era
by: Wilkins, Grant, et al.
Published: (2026) -
EcoServe: Designing Carbon-Aware AI Inference Systems
by: Li, Yueying, et al.
Published: (2025) -
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024) -
Towards Resource-Efficient Compound AI Systems
by: Chaudhry, Gohar Irfan, et al.
Published: (2025)