Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
Fuente:
arXiv
Saved in:
| Main Authors: | Delavande, Julien, Pierrard, Regis, Luccioni, Sasha |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Small Talk, Big Impact: The Energy Cost of Thanking AI
by: Delavande, Julien, et al.
Published: (2026)
by: Delavande, Julien, et al.
Published: (2026)
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
by: Delavande, Julien, et al.
Published: (2025)
by: Delavande, Julien, et al.
Published: (2025)
Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines
by: Lambert, Katherine, et al.
Published: (2026)
by: Lambert, Katherine, et al.
Published: (2026)
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025)
by: Fernandez, Jared, et al.
Published: (2025)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
Energy and Carbon Considerations of Fine-Tuning BERT
by: Wang, Xiaorong, et al.
Published: (2023)
by: Wang, Xiaorong, et al.
Published: (2023)
Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
by: LaCroix, Travis, et al.
Published: (2026)
by: LaCroix, Travis, et al.
Published: (2026)
A Paired Testing Protocol for Batch-Conditioned Refusal Robustness in LLM Serving
by: Kadadekar, Sahil
Published: (2026)
by: Kadadekar, Sahil
Published: (2026)
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
by: Zhao, Yilong, et al.
Published: (2023)
by: Zhao, Yilong, et al.
Published: (2023)
SMDP-Based Dynamic Batching for Improving Responsiveness and Energy Efficiency of Batch Services
by: Xu, Yaodan, et al.
Published: (2025)
by: Xu, Yaodan, et al.
Published: (2025)
RW-TTT: Batched Serving for Request-Owned Test-Time Training State
by: Yang, Jian, et al.
Published: (2026)
by: Yang, Jian, et al.
Published: (2026)
Oaken: Fast and Efficient LLM Serving with Online-Offline Hybrid KV Cache Quantization
by: Kim, Minsu, et al.
Published: (2025)
by: Kim, Minsu, et al.
Published: (2025)
SAW-INT4: System-Aware 4-Bit KV-Cache Quantization for Real-World LLM Serving
by: Jia, Jinda, et al.
Published: (2026)
by: Jia, Jinda, et al.
Published: (2026)
MorphServe: Efficient and Workload-Aware LLM Serving via Runtime Quantized Layer Swapping and KV Cache Resizing
by: Su, Zhaoyuan, et al.
Published: (2025)
by: Su, Zhaoyuan, et al.
Published: (2025)
LLM-PQ: Serving LLM on Heterogeneous Clusters with Phase-Aware Partition and Adaptive Quantization
by: Zhao, Juntao, et al.
Published: (2024)
by: Zhao, Juntao, et al.
Published: (2024)
BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching
by: Zhao, Yilong, et al.
Published: (2024)
by: Zhao, Yilong, et al.
Published: (2024)
FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
CliqueParcel: An Approach For Batching LLM Prompts That Jointly Optimizes Efficiency And Faithfulness
by: Liu, Jiayi, et al.
Published: (2024)
by: Liu, Jiayi, et al.
Published: (2024)
Quantamination: Dynamic Quantization Leaks Your Data Across the Batch
by: Foerster, Hanna, et al.
Published: (2026)
by: Foerster, Hanna, et al.
Published: (2026)
Symphony: Optimized DNN Model Serving using Deferred Batch Scheduling
by: Chen, Lequn, et al.
Published: (2023)
by: Chen, Lequn, et al.
Published: (2023)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Batched Energy-Entropy acquisition for Bayesian Optimization
by: Teufel, Felix, et al.
Published: (2024)
by: Teufel, Felix, et al.
Published: (2024)
SUN: Shared Use of Next-token Prediction for Efficient Multi-LLM Disaggregated Serving
by: Woo, Sunghyeon, et al.
Published: (2026)
by: Woo, Sunghyeon, et al.
Published: (2026)
Rethinking Efficiency in Neural Combinatorial Optimization: Batched Preference Optimization with Mamba
by: Xu, Zhenxing, et al.
Published: (2026)
by: Xu, Zhenxing, et al.
Published: (2026)
DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU Multiplexing
by: Gao, Lei, et al.
Published: (2025)
by: Gao, Lei, et al.
Published: (2025)
Apt-Serve: Adaptive Request Scheduling on Hybrid Cache for Scalable LLM Inference Serving
by: Gao, Shihong, et al.
Published: (2025)
by: Gao, Shihong, et al.
Published: (2025)
EnergyLens: Interpretable Closed-Form Energy Models for Multimodal LLM Inference Serving
by: Palladino, Vittorio, et al.
Published: (2026)
by: Palladino, Vittorio, et al.
Published: (2026)
Position: Key Claims in LLM Research Have a Long Tail of Footnotes
by: Rogers, Anna, et al.
Published: (2023)
by: Rogers, Anna, et al.
Published: (2023)
GreenServ: Energy-Efficient Context-Aware Dynamic Routing for Multi-Model LLM Inference
by: Ziller, Thomas, et al.
Published: (2026)
by: Ziller, Thomas, et al.
Published: (2026)
Operationalizing Quantized Disentanglement
by: Barin-Pacela, Vitoria, et al.
Published: (2025)
by: Barin-Pacela, Vitoria, et al.
Published: (2025)
Towards Understanding Neural Collapse: The Effects of Batch Normalization and Weight Decay
by: Pan, Leyan, et al.
Published: (2023)
by: Pan, Leyan, et al.
Published: (2023)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
by: Tian, Jian, et al.
Published: (2025)
by: Tian, Jian, et al.
Published: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
by: Husom, Erik Johannes, et al.
Published: (2025)
by: Husom, Erik Johannes, et al.
Published: (2025)
Understanding Behavior Cloning with Action Quantization
by: Cao, Haoqun, et al.
Published: (2026)
by: Cao, Haoqun, et al.
Published: (2026)
CrossQ: Batch Normalization in Deep Reinforcement Learning for Greater Sample Efficiency and Simplicity
by: Bhatt, Aditya, et al.
Published: (2019)
by: Bhatt, Aditya, et al.
Published: (2019)
The CAP Principle for LLM Serving: A Survey of Long-Context Large Language Model Serving
by: Zeng, Pai, et al.
Published: (2024)
by: Zeng, Pai, et al.
Published: (2024)
How Particle-System Random Batch Methods Enhance Graph Transformer: Memory Efficiency and Parallel Computing Strategy
by: Liu, Hanwen, et al.
Published: (2025)
by: Liu, Hanwen, et al.
Published: (2025)
Understanding Quantization of Optimizer States in LLM Pre-training: Dynamics of State Staleness and Effectiveness of State Resets
by: Topollai, Kristi, et al.
Published: (2026)
by: Topollai, Kristi, et al.
Published: (2026)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
MEPIC: Memory Efficient Position Independent Caching for LLM Serving
by: Wang, Qian, et al.
Published: (2025)
by: Wang, Qian, et al.
Published: (2025)
Similar Items
-
Small Talk, Big Impact: The Energy Cost of Thanking AI
by: Delavande, Julien, et al.
Published: (2026) -
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
by: Delavande, Julien, et al.
Published: (2025) -
Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines
by: Lambert, Katherine, et al.
Published: (2026) -
Energy Considerations of Large Language Model Inference and Efficiency Optimizations
by: Fernandez, Jared, et al.
Published: (2025) -
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)