SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference
Fuente:
arXiv
Guardado en:
| Autores principales: | Cavagna, Hiari Pizzini, Proia, Andrea, Madella, Giacomo, Esposito, Giovanni B., Antici, Francesco, Cesarini, Daniele, Kiziltan, Zeynep, Bartolini, Andrea |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2025)
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2025)
From Data Center IoT Telemetry to Data Analytics Chatbots -- Virtual Knowledge Graph is All You Need
por: Khan, Junaid Ahmed, et al.
Publicado: (2025)
por: Khan, Junaid Ahmed, et al.
Publicado: (2025)
Modeling and Controlling Many-Core HPC Processors: an Alternative to PID and Moving Average Algorithms
por: Bambini, Giovanni, et al.
Publicado: (2024)
por: Bambini, Giovanni, et al.
Publicado: (2024)
Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-Chips
por: Parisi, Emanuele, et al.
Publicado: (2024)
por: Parisi, Emanuele, et al.
Publicado: (2024)
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
por: Chang, Dali, et al.
Publicado: (2026)
por: Chang, Dali, et al.
Publicado: (2026)
On Combining Two Server Control Policies for Energy Efficiency
por: Dai, Jingze, et al.
Publicado: (2025)
por: Dai, Jingze, et al.
Publicado: (2025)
Silicon Showdown: Performance, Efficiency, and Ecosystem Barriers in Consumer-Grade LLM Inference
por: Javat, Abdurrahman, et al.
Publicado: (2026)
por: Javat, Abdurrahman, et al.
Publicado: (2026)
Statistical Modeling and Uncertainty Estimation of LLM Inference Systems
por: Ray, Kaustabha, et al.
Publicado: (2025)
por: Ray, Kaustabha, et al.
Publicado: (2025)
Monte Cimone v2: Down the Road of RISC-V High-Performance Computers
por: Venieri, Emanuele, et al.
Publicado: (2025)
por: Venieri, Emanuele, et al.
Publicado: (2025)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
por: Ziller, Thomas, et al.
Publicado: (2026)
por: Ziller, Thomas, et al.
Publicado: (2026)
Forecasting LLM Inference Performance via Hardware-Agnostic Analytical Modeling
por: Patwari, Rajeev, et al.
Publicado: (2025)
por: Patwari, Rajeev, et al.
Publicado: (2025)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
por: Wang, Qiang, et al.
Publicado: (2024)
por: Wang, Qiang, et al.
Publicado: (2024)
Characterize LSM-tree Compaction Performance via On-Device LLM Inference
por: Ding, Jiabiao, et al.
Publicado: (2026)
por: Ding, Jiabiao, et al.
Publicado: (2026)
Fine-Grained Energy Prediction For Parallellized LLM Inference With PIE-P
por: Dutt, Anurag, et al.
Publicado: (2025)
por: Dutt, Anurag, et al.
Publicado: (2025)
GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM Serving
por: Liu, Qunyou, et al.
Publicado: (2025)
por: Liu, Qunyou, et al.
Publicado: (2025)
SparseInfer: Training-free Prediction of Activation Sparsity for Fast LLM Inference
por: Shin, Jiho, et al.
Publicado: (2024)
por: Shin, Jiho, et al.
Publicado: (2024)
Energy Efficiency Analysis of Active RIS-enhanced Wireless Network under Power-Sum Constraint
por: Xin, Jingdie, et al.
Publicado: (2025)
por: Xin, Jingdie, et al.
Publicado: (2025)
It's Not Easy Being Green: On the Energy Efficiency of Programming Languages
por: van Kempen, Nicolas, et al.
Publicado: (2024)
por: van Kempen, Nicolas, et al.
Publicado: (2024)
ETM2: Empowering Traditional Memory Bandwidth Regulation using ETM
por: Zuepke, Alexander, et al.
Publicado: (2026)
por: Zuepke, Alexander, et al.
Publicado: (2026)
How to Increase Energy Efficiency with a Single Linux Command
por: Jelvani, Alborz, et al.
Publicado: (2025)
por: Jelvani, Alborz, et al.
Publicado: (2025)
An Inquiry into Datacenter TCO for LLM Inference with FP8
por: Kim, Jiwoo, et al.
Publicado: (2025)
por: Kim, Jiwoo, et al.
Publicado: (2025)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
por: Fu, Zizhuo, et al.
Publicado: (2025)
por: Fu, Zizhuo, et al.
Publicado: (2025)
Greener Deep Reinforcement Learning: Analysis of Energy and Carbon Efficiency Across Atari Benchmarks
por: Gardner, Jason, et al.
Publicado: (2025)
por: Gardner, Jason, et al.
Publicado: (2025)
CARINA: Carbon-Aware Execution of Recurrent Industrial Analytics
por: Farooq, Muhammad Umar
Publicado: (2026)
por: Farooq, Muhammad Umar
Publicado: (2026)
PARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation
por: An, Zihao, et al.
Publicado: (2025)
por: An, Zihao, et al.
Publicado: (2025)
MarginGate: Sparse Margin-Triggered Verification for Batch-Invariant LLM Inference
por: Chu, Kexin, et al.
Publicado: (2026)
por: Chu, Kexin, et al.
Publicado: (2026)
Faster LLM Inference using DBMS-Inspired Preemption and Cache Replacement Policies
por: Kim, Kyoungmin, et al.
Publicado: (2024)
por: Kim, Kyoungmin, et al.
Publicado: (2024)
Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
por: Zaeed, Mohammad, et al.
Publicado: (2025)
por: Zaeed, Mohammad, et al.
Publicado: (2025)
Computational Algorithms for the Product Form Solution of Closed Queuing Networks with Finite Buffers and Skip-Over Policy
por: Balbo, Gianfranco, et al.
Publicado: (2024)
por: Balbo, Gianfranco, et al.
Publicado: (2024)
Enabling Heterogeneous Performance Analysis for Scientific Workloads
por: Graczyk, Maksymilian, et al.
Publicado: (2025)
por: Graczyk, Maksymilian, et al.
Publicado: (2025)
5G Cellular -- An Energy Efficiency Perspective
por: Panchal, Deven
Publicado: (2024)
por: Panchal, Deven
Publicado: (2024)
Performance Characterization of Expert Router for Scalable LLM Inference
por: Pichlmeier, Josef, et al.
Publicado: (2024)
por: Pichlmeier, Josef, et al.
Publicado: (2024)
Dissecting Embedding Bag Performance in DLRM Inference
por: Ambati, Chandrish, et al.
Publicado: (2025)
por: Ambati, Chandrish, et al.
Publicado: (2025)
The Multiserver-Job Stochastic Recurrence Equation for Cloud Computing Performance Evaluation
por: Baccelli, Francois, et al.
Publicado: (2026)
por: Baccelli, Francois, et al.
Publicado: (2026)
RAPID-LLM: Resilience-Aware Performance analysis of Infrastructure for Distributed LLM Training and Inference
por: Karfakis, George, et al.
Publicado: (2025)
por: Karfakis, George, et al.
Publicado: (2025)
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
por: Rodrigo, Javier J. Poveda, et al.
Publicado: (2025)
por: Rodrigo, Javier J. Poveda, et al.
Publicado: (2025)
Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs
por: Georganas, Evangelos, et al.
Publicado: (2025)
por: Georganas, Evangelos, et al.
Publicado: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
por: Zhang, Li, et al.
Publicado: (2025)
por: Zhang, Li, et al.
Publicado: (2025)
Analytic Roofline Modeling and Energy Analysis of LULESH Proxy Application on Multi-Core Clusters
por: Afzal, Ayesha, et al.
Publicado: (2024)
por: Afzal, Ayesha, et al.
Publicado: (2024)
Towards Multi-dimensional Elasticity for Pervasive Stream Processing Services
por: Sedlak, Boris, et al.
Publicado: (2025)
por: Sedlak, Boris, et al.
Publicado: (2025)
Ejemplares similares
-
Assessing Tenstorrent's RISC-V MatMul Acceleration Capabilities
por: Cavagna, Hiari Pizzini, et al.
Publicado: (2025) -
From Data Center IoT Telemetry to Data Analytics Chatbots -- Virtual Knowledge Graph is All You Need
por: Khan, Junaid Ahmed, et al.
Publicado: (2025) -
Modeling and Controlling Many-Core HPC Processors: an Alternative to PID and Moving Average Algorithms
por: Bambini, Giovanni, et al.
Publicado: (2024) -
Assessing the Performance of OpenTitan as Cryptographic Accelerator in Secure Open-Hardware System-on-Chips
por: Parisi, Emanuele, et al.
Publicado: (2024) -
Shifting the Sweet Spot: High-Performance Matrix-Free Method for High-Order Elasticity
por: Chang, Dali, et al.
Publicado: (2026)