Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Stojkovic, Jovan, Choukse, Esha, Zhang, Chaojie, Goiri, Inigo, Torrellas, Josep |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024)
by: Stojkovic, Jovan, et al.
Published: (2024)
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023)
by: Patel, Pratyush, et al.
Published: (2023)
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
by: Iliakopoulou, Nikoleta, et al.
Published: (2024)
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)
by: Stojkovic, Jovan, et al.
Published: (2025)
Exploring the Efficiency of 3D-Stacked AI Chip Architecture for LLM Inference with Voxel
by: Liu, Yiqi, et al.
Published: (2026)
by: Liu, Yiqi, et al.
Published: (2026)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
Understanding Bottlenecks for Efficiently Serving LLM Inference With KV Offloading
by: Meng, William, et al.
Published: (2025)
by: Meng, William, et al.
Published: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
by: Zhang, Chen, et al.
Published: (2026)
by: Zhang, Chen, et al.
Published: (2026)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
by: Li, Jiamin, et al.
Published: (2025)
by: Li, Jiamin, et al.
Published: (2025)
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache
by: Lin, Bin, et al.
Published: (2024)
by: Lin, Bin, et al.
Published: (2024)
Improving Multi-Instance GPU Efficiency via Sub-Entry Sharing TLB Design
by: Li, Bingyao, et al.
Published: (2024)
by: Li, Bingyao, et al.
Published: (2024)
Automated Deep Neural Network Inference Partitioning for Distributed Embedded Systems
by: Kreß, Fabian, et al.
Published: (2024)
by: Kreß, Fabian, et al.
Published: (2024)
HgPCN: A Heterogeneous Architecture for E2E Embedded Point Cloud Inference
by: Gao, Yiming, et al.
Published: (2025)
by: Gao, Yiming, et al.
Published: (2025)
DUET: Disaggregated Hybrid Mamba-Transformer LLMs with Prefill and Decode-Specific Packages
by: Kanani, Alish, et al.
Published: (2026)
by: Kanani, Alish, et al.
Published: (2026)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
by: Xu, Weihong, et al.
Published: (2025)
by: Xu, Weihong, et al.
Published: (2025)
Adaptive KV Cache Reuse for Fast Long-Context LLM Serving
by: li, Fei, et al.
Published: (2026)
by: li, Fei, et al.
Published: (2026)
Adaptive Multi-Objective Tiered Storage Configuration for KV Cache in LLM Service
by: Zheng, Xianzhe, et al.
Published: (2026)
by: Zheng, Xianzhe, et al.
Published: (2026)
Fine-Grained Power and Energy Attribution on AMD GPU/APU-Based Exascale Nodes
by: McDaniel, Adam, et al.
Published: (2026)
by: McDaniel, Adam, et al.
Published: (2026)
DiP: A Scalable, Energy-Efficient Systolic Array for Matrix Multiplication Acceleration
by: Abdelmaksoud, Ahmed J., et al.
Published: (2024)
by: Abdelmaksoud, Ahmed J., et al.
Published: (2024)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
by: Kubo, Tatsuya, et al.
Published: (2025)
by: Kubo, Tatsuya, et al.
Published: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
by: Liu, Lian, et al.
Published: (2026)
by: Liu, Lian, et al.
Published: (2026)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
by: Chu, Xiaoyu, et al.
Published: (2024)
by: Chu, Xiaoyu, et al.
Published: (2024)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
by: Kurzynski, Marco, et al.
Published: (2025)
by: Kurzynski, Marco, et al.
Published: (2025)
GreenLLM: Disaggregating Large Language Model Serving on Heterogeneous GPUs for Lower Carbon Emissions
by: Shi, Tianyao, et al.
Published: (2024)
by: Shi, Tianyao, et al.
Published: (2024)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
by: Kwak, Hyunseok, et al.
Published: (2025)
by: Kwak, Hyunseok, et al.
Published: (2025)
A Scalable NorthPole System with End-to-End Vertical Integration for Low-Latency and Energy-Efficient LLM Inference
by: DeBole, Michael V., et al.
Published: (2025)
by: DeBole, Michael V., et al.
Published: (2025)
SwarmIO: Towards 100 Million IOPS SSD Emulation for Next-generation GPU-centric Storage Systems
by: Kim, Hyeseong, et al.
Published: (2026)
by: Kim, Hyeseong, et al.
Published: (2026)
MIMDRAM: An End-to-End Processing-Using-DRAM System for High-Throughput, Energy-Efficient and Programmer-Transparent Multiple-Instruction Multiple-Data Processing
by: Oliveira, Geraldo F., et al.
Published: (2024)
by: Oliveira, Geraldo F., et al.
Published: (2024)
SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference
by: Zhang, Hengrui, et al.
Published: (2025)
by: Zhang, Hengrui, et al.
Published: (2025)
Fusing Depthwise and Pointwise Convolutions for Efficient Inference on GPUs
by: Qararyah, Fareed, et al.
Published: (2024)
by: Qararyah, Fareed, et al.
Published: (2024)
A High Energy-Efficiency Multi-core Neuromorphic Architecture for Deep SNN Training
by: Li, Mingjing, et al.
Published: (2024)
by: Li, Mingjing, et al.
Published: (2024)
Profiling Concurrent Vision Inference Workloads on NVIDIA Jetson -- Extended
by: Chakraborty, Abhinaba, et al.
Published: (2025)
by: Chakraborty, Abhinaba, et al.
Published: (2025)
EPAC: The Last Dance
by: Mantovani, Filippo, et al.
Published: (2026)
by: Mantovani, Filippo, et al.
Published: (2026)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
by: Tschand, Arya, et al.
Published: (2024)
by: Tschand, Arya, et al.
Published: (2024)
Insights into DeepSeek-V3: Scaling Challenges and Reflections on Hardware for AI Architectures
by: Zhao, Chenggang, et al.
Published: (2025)
by: Zhao, Chenggang, et al.
Published: (2025)
PIM-AI: A Novel Architecture for High-Efficiency LLM Inference
by: Ortega, Cristobal, et al.
Published: (2024)
by: Ortega, Cristobal, et al.
Published: (2024)
MegIS: High-Performance, Energy-Efficient, and Low-Cost Metagenomic Analysis with In-Storage Processing
by: Ghiasi, Nika Mansouri, et al.
Published: (2024)
by: Ghiasi, Nika Mansouri, et al.
Published: (2024)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
by: Kreuzer, Anke, et al.
Published: (2019)
by: Kreuzer, Anke, et al.
Published: (2019)
Similar Items
-
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
by: Stojkovic, Jovan, et al.
Published: (2024) -
Splitwise: Efficient generative LLM inference using phase splitting
by: Patel, Pratyush, et al.
Published: (2023) -
Rearchitecting Datacenter Lifecycle for AI: A TCO-Driven Framework
by: Stojkovic, Jovan, et al.
Published: (2025) -
Chameleon: Adaptive Caching and Scheduling for Many-Adapter LLM Inference Environments
by: Iliakopoulou, Nikoleta, et al.
Published: (2024) -
TAPAS: Thermal- and Power-Aware Scheduling for LLM Inference in Cloud Platforms
by: Stojkovic, Jovan, et al.
Published: (2025)