Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Maliakel, Paul Joe, Ilager, Shashikant, Brandic, Ivona |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FLIGAN: Enhancing Federated Learning with Incomplete Data using GAN
by: Maliakel, Paul Joe, et al.
Published: (2024)
by: Maliakel, Paul Joe, et al.
Published: (2024)
ABBA-VSM: Time Series Classification using Symbolic Representation on the Edge
by: Kanatbekova, Meerzhan, et al.
Published: (2024)
by: Kanatbekova, Meerzhan, et al.
Published: (2024)
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
by: May, Daniel, et al.
Published: (2024)
by: May, Daniel, et al.
Published: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026)
by: Šabanović, Ahmed, et al.
Published: (2026)
GREEN-CODE: Learning to Optimize Energy Efficiency in LLM-based Code Generation
by: Ilager, Shashikant, et al.
Published: (2025)
by: Ilager, Shashikant, et al.
Published: (2025)
A Decentralized and Self-Adaptive Approach for Monitoring Volatile Edge Environments
by: Ilager, Shashikant, et al.
Published: (2024)
by: Ilager, Shashikant, et al.
Published: (2024)
FRESCO: Fast and Reliable Edge Offloading with Reputation-based Hybrid Smart Contracts
by: Zilic, Josip, et al.
Published: (2024)
by: Zilic, Josip, et al.
Published: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
by: Chu, Xiaoyu, et al.
Published: (2024)
by: Chu, Xiaoyu, et al.
Published: (2024)
Limits of quantum generative models with classical sampling hardness
by: Herbst, Sabrina, et al.
Published: (2025)
by: Herbst, Sabrina, et al.
Published: (2025)
On Optimizing Hyperparameters for Quantum Neural Networks
by: Herbst, Sabrina, et al.
Published: (2024)
by: Herbst, Sabrina, et al.
Published: (2024)
Streaming IoT Data and the Quantum Edge: A Classic/Quantum Machine Learning Use Case
by: Herbst, Sabrina, et al.
Published: (2024)
by: Herbst, Sabrina, et al.
Published: (2024)
Clustered Federated Learning with Hierarchical Knowledge Distillation
by: Ahmad, Sabtain, et al.
Published: (2025)
by: Ahmad, Sabtain, et al.
Published: (2025)
Exploring Channel Distinguishability in Local Neighborhoods of the Model Space in Quantum Neural Networks
by: Herbst, Sabrina, et al.
Published: (2024)
by: Herbst, Sabrina, et al.
Published: (2024)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
by: Song, Zhiye, et al.
Published: (2026)
by: Song, Zhiye, et al.
Published: (2026)
LLM Zeroth-Order Fine-Tuning is an Inference Workload
by: Li, Zelin, et al.
Published: (2026)
by: Li, Zelin, et al.
Published: (2026)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
Bit-Exact AI Inference Verification Without Performance Tradeoffs
by: Cankaya, Naci
Published: (2026)
by: Cankaya, Naci
Published: (2026)
Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve
by: Agrawal, Amey, et al.
Published: (2024)
by: Agrawal, Amey, et al.
Published: (2024)
Confidential LLM Inference: Performance and Cost Across CPU and GPU TEEs
by: Chrapek, Marcin, et al.
Published: (2025)
by: Chrapek, Marcin, et al.
Published: (2025)
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
by: Gopal, Nikhil, et al.
Published: (2026)
by: Gopal, Nikhil, et al.
Published: (2026)
Characterizing WebGPU Dispatch Overhead for LLM Inference Across Four GPU Vendors, Three Backends, and Three Browsers
by: Maczan, Jędrzej
Published: (2026)
by: Maczan, Jędrzej
Published: (2026)
Forecasting GPU Performance for Deep Learning Training and Inference
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Improving Predictions on Highly Unbalanced Data Using Open Source Synthetic Data Upsampling
by: Krchova, Ivona, et al.
Published: (2025)
by: Krchova, Ivona, et al.
Published: (2025)
Navigating the Exploration-Exploitation Tradeoff in Inference-Time Scaling of Diffusion Models
by: Su, Xun, et al.
Published: (2025)
by: Su, Xun, et al.
Published: (2025)
Scaling On-Device GPU Inference for Large Generative Models
by: Tang, Jiuqiang, et al.
Published: (2025)
by: Tang, Jiuqiang, et al.
Published: (2025)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
by: Deng, Weishu, et al.
Published: (2025)
by: Deng, Weishu, et al.
Published: (2025)
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
by: Ye, Zicong, et al.
Published: (2025)
by: Ye, Zicong, et al.
Published: (2025)
A Framework for Carbon-aware Real-Time Workload Management in Clouds using Renewables-driven Cores
by: Hewage, Tharindu B., et al.
Published: (2024)
by: Hewage, Tharindu B., et al.
Published: (2024)
Characterizing and Optimizing LLM Inference Workloads on CPU-GPU Coupled Architectures
by: Vellaisamy, Prabhu, et al.
Published: (2025)
by: Vellaisamy, Prabhu, et al.
Published: (2025)
Aging-aware CPU Core Management for Embodied Carbon Amortization in Cloud LLM Inference
by: Hewage, Tharindu B., et al.
Published: (2025)
by: Hewage, Tharindu B., et al.
Published: (2025)
SpecOffload: Unlocking Latent GPU Capacity for LLM Inference on Resource-Constrained Devices
by: Zhuge, Xiangwen, et al.
Published: (2025)
by: Zhuge, Xiangwen, et al.
Published: (2025)
Green Prompting: Characterizing Prompt-driven Energy Costs of LLM Inference
by: Adamska, Marta, et al.
Published: (2025)
by: Adamska, Marta, et al.
Published: (2025)
LLM-Pilot: Characterize and Optimize Performance of your LLM Inference Services
by: Łazuka, Małgorzata, et al.
Published: (2024)
by: Łazuka, Małgorzata, et al.
Published: (2024)
LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
by: Tummalapalli, Pranay, et al.
Published: (2026)
by: Tummalapalli, Pranay, et al.
Published: (2026)
Accelerating Sparse Transformer Inference on GPU
by: Dai, Wenhao, et al.
Published: (2025)
by: Dai, Wenhao, et al.
Published: (2025)
Understanding the Performance Horizon of the Latest ML Workloads with NonGEMM Workloads
by: Karami, Rachid, et al.
Published: (2024)
by: Karami, Rachid, et al.
Published: (2024)
Llamas on the Web: Memory-Efficient, Performance-Portable, and Multi-Precision LLM Inference with WebGPU
by: Levine, Reese, et al.
Published: (2026)
by: Levine, Reese, et al.
Published: (2026)
On the Effect of Sampling Diversity in Scaling LLM Inference
by: Wang, Tianchun, et al.
Published: (2025)
by: Wang, Tianchun, et al.
Published: (2025)
Similar Items
-
FLIGAN: Enhancing Federated Learning with Incomplete Data using GAN
by: Maliakel, Paul Joe, et al.
Published: (2024) -
ABBA-VSM: Time Series Classification using Symbolic Representation on the Edge
by: Kanatbekova, Meerzhan, et al.
Published: (2024) -
DynaSplit: A Hardware-Software Co-Design Framework for Energy-Aware Inference on Edge
by: May, Daniel, et al.
Published: (2024) -
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026) -
INAR-VL: Input-Aware Routing for Edge-Cloud Vision-Language Inference
by: Šabanović, Ahmed, et al.
Published: (2026)