EnergyLens: Predictive Energy-Aware Exploration for Multi-GPU LLM Inference Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Zhiye, Lee, Kyungmi, Lee, Eun Kyung, Zhang, Xin, Eilam, Tamar, Chandrakasan, Anantha P. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026)
by: Lee, Kyungmi, et al.
Published: (2026)
EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving
by: Palladino, Vittorio, et al.
Published: (2026)
by: Palladino, Vittorio, et al.
Published: (2026)
LEDRO: LLM-Enhanced Design Space Reduction and Optimization for Analog Circuits
by: Kochar, Dimple Vijay, et al.
Published: (2024)
by: Kochar, Dimple Vijay, et al.
Published: (2024)
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
by: Choochotkaew, Sunyanan, et al.
Published: (2024)
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)
by: Recasens, Pol G., et al.
Published: (2025)
FedCore: Straggler-Free Federated Learning with Distributed Coresets
by: Guo, Hongpeng, et al.
Published: (2024)
by: Guo, Hongpeng, et al.
Published: (2024)
Characterizing LLM Inference Energy-Performance Tradeoffs across Workloads and GPU Scaling
by: Maliakel, Paul Joe, et al.
Published: (2025)
by: Maliakel, Paul Joe, et al.
Published: (2025)
Online GPU Energy Optimization with Switching-Aware Bandits
by: Xu, Xiongxiao, et al.
Published: (2024)
by: Xu, Xiongxiao, et al.
Published: (2024)
SLO-aware GPU Frequency Scaling for Energy Efficient LLM Inference Serving
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
by: Kakolyris, Andreas Kosmas, et al.
Published: (2024)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
Multi-LLM Adaptive Conformal Inference for Reliable LLM Responses
by: Noh, Kangjun, et al.
Published: (2026)
by: Noh, Kangjun, et al.
Published: (2026)
AGFT: An Adaptive GPU Frequency Tuner for Real-Time LLM Inference Optimization
by: Ye, Zicong, et al.
Published: (2025)
by: Ye, Zicong, et al.
Published: (2025)
Adaptive Block-Scaled Data Types
by: Cook, Jack, et al.
Published: (2026)
by: Cook, Jack, et al.
Published: (2026)
IRONIC: Coherence-Aware Reasoning Chains for Multi-Modal Sarcasm Detection
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
LLM-Guided Runtime Parameter Optimization for Energy-Efficient Model Inference
by: Crumpacker, Katelyn, et al.
Published: (2026)
by: Crumpacker, Katelyn, et al.
Published: (2026)
Forecasting GPU Performance for Deep Learning Training and Inference
by: Lee, Seonho, et al.
Published: (2024)
by: Lee, Seonho, et al.
Published: (2024)
Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding
by: Park, Jihoon, et al.
Published: (2025)
by: Park, Jihoon, et al.
Published: (2025)
Semantic-Aware Gaussian Process Calibration with Structured Layerwise Kernels for Deep Neural Networks
by: Lee, Kyung-hwan, et al.
Published: (2025)
by: Lee, Kyung-hwan, et al.
Published: (2025)
Energy-Aware Dynamic Neural Inference
by: Bullo, Marcello, et al.
Published: (2024)
by: Bullo, Marcello, et al.
Published: (2024)
Energy-Aware DNN Graph Optimization
by: Wang, Yu, et al.
Published: (2020)
by: Wang, Yu, et al.
Published: (2020)
SparseDVFS: Sparse-Aware DVFS for Energy-Efficient Edge Inference
by: Zhang, Ziyang, et al.
Published: (2026)
by: Zhang, Ziyang, et al.
Published: (2026)
HGCA: Hybrid GPU-CPU Attention for Long Context LLM Inference
by: Deng, Weishu, et al.
Published: (2025)
by: Deng, Weishu, et al.
Published: (2025)
An Enhanced Projection Pursuit Tree Classifier with Visual Methods for Assessing Algorithmic Improvements
by: da Silva, Natalia, et al.
Published: (2026)
by: da Silva, Natalia, et al.
Published: (2026)
Interactive Graphics for Visually Diagnosing Forest Classifiers in R
by: da Silva, Natalia, et al.
Published: (2017)
by: da Silva, Natalia, et al.
Published: (2017)
Watt Counts: Energy-Aware Benchmark for Sustainable LLM Inference on Heterogeneous GPU Architectures
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
by: Argerich, Mauricio Fadel, et al.
Published: (2026)
Correctness-Optimized Residual Activation Lens (CORAL): Transferrable and Calibration-Aware Inference-Time Steering
by: Miao, Miranda Muqing, et al.
Published: (2026)
by: Miao, Miranda Muqing, et al.
Published: (2026)
NanoFlux: Adversarial Dual-LLM Evaluation and Distillation For Multi-Domain Reasoning
by: Anantha, Raviteja, et al.
Published: (2025)
by: Anantha, Raviteja, et al.
Published: (2025)
Learning Short-Term and Long-Term Patterns of High-Order Dynamics in Real-World Networks
by: Ko, Yunyong, et al.
Published: (2025)
by: Ko, Yunyong, et al.
Published: (2025)
Harvest: Opportunistic Peer-to-Peer GPU Caching for LLM Inference
by: Gopal, Nikhil, et al.
Published: (2026)
by: Gopal, Nikhil, et al.
Published: (2026)
ModeSwitch-LLM: A Lightweight Phase-Aware Controller for Cross-Mode LLM Inference on a Single GPU
by: Sunesh, Aman, et al.
Published: (2026)
by: Sunesh, Aman, et al.
Published: (2026)
LIDS: LLM Summary Inference Under the Layered Lens
by: Park, Dylan, et al.
Published: (2026)
by: Park, Dylan, et al.
Published: (2026)
Transformers as Intrinsic Optimizers: Forward Inference through the Energy Principle
by: Ren, Ruifeng, et al.
Published: (2025)
by: Ren, Ruifeng, et al.
Published: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
by: Zhang, Haolin, et al.
Published: (2025)
by: Zhang, Haolin, et al.
Published: (2025)
From Prompts to Power: Measuring the Energy Footprint of LLM Inference
by: Caravaca, Francisco, et al.
Published: (2025)
by: Caravaca, Francisco, et al.
Published: (2025)
PULSE-ICU: A Pretrained Unified Long-Sequence Encoder for Multi-task Prediction in Intensive Care Units
by: Jang, Sejeong, et al.
Published: (2025)
by: Jang, Sejeong, et al.
Published: (2025)
Variational Inference Optimized Using the Curved Geometry of Coupled Free Energy
by: Nelson, Kenric, et al.
Published: (2025)
by: Nelson, Kenric, et al.
Published: (2025)
Scaling On-Device GPU Inference for Large Generative Models
by: Tang, Jiuqiang, et al.
Published: (2025)
by: Tang, Jiuqiang, et al.
Published: (2025)
New User Event Prediction Through the Lens of Causal Inference
by: Yuchi, Henry Shaowu, et al.
Published: (2024)
by: Yuchi, Henry Shaowu, et al.
Published: (2024)
Efficient Edge LLMs Deployment via HessianAware Quantization and CPU GPU Collaborative
by: Zhang, Tuo, et al.
Published: (2025)
by: Zhang, Tuo, et al.
Published: (2025)
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025)
by: Fernandez, Jared, et al.
Published: (2025)
Similar Items
-
EnergAIzer: Fast and Accurate GPU Power Estimation Framework for AI Workloads
by: Lee, Kyungmi, et al.
Published: (2026) -
EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving
by: Palladino, Vittorio, et al.
Published: (2026) -
LEDRO: LLM-Enhanced Design Space Reduction and Optimization for Analog Circuits
by: Kochar, Dimple Vijay, et al.
Published: (2024) -
A Robust Power Model Training Framework for Cloud Native Runtime Energy Metric Exporter
by: Choochotkaew, Sunyanan, et al.
Published: (2024) -
Mind the Memory Gap: Unveiling GPU Bottlenecks in Large-Batch LLM Inference
by: Recasens, Pol G., et al.
Published: (2025)