Energy Considerations of Large Language Model Inference and Efficiency Optimizations
Fuente:
arXiv
Saved in:
| Main Authors: | Fernandez, Jared, Na, Clara, Tiwari, Vashisth, Bisk, Yonatan, Luccioni, Sasha, Strubell, Emma |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Energy and Carbon Considerations of Fine-Tuning BERT
by: Wang, Xiaorong, et al.
Published: (2023)
by: Wang, Xiaorong, et al.
Published: (2023)
Gradient Localization Improves Lifelong Pretraining of Language Models
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024)
by: Chang, Yingshan, et al.
Published: (2024)
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
by: Luccioni, Alexandra Sasha, et al.
Published: (2023)
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)
by: Fernandez, Jared, et al.
Published: (2024)
Scalable Data Ablation Approximations for Language Models through Modular Training and Merging
by: Na, Clara, et al.
Published: (2024)
by: Na, Clara, et al.
Published: (2024)
Holistically Evaluating the Environmental Impact of Creating Language Models
by: Morrison, Jacob, et al.
Published: (2025)
by: Morrison, Jacob, et al.
Published: (2025)
From Efficiency Gains to Rebound Effects: The Problem of Jevons' Paradox in AI's Polarized Environmental Debate
by: Luccioni, Alexandra Sasha, et al.
Published: (2025)
by: Luccioni, Alexandra Sasha, et al.
Published: (2025)
Understanding Efficiency: Quantization, Batching, and Serving Strategies in LLM Energy Use
by: Delavande, Julien, et al.
Published: (2026)
by: Delavande, Julien, et al.
Published: (2026)
Looking beyond the next token
by: Thankaraj, Abitha, et al.
Published: (2025)
by: Thankaraj, Abitha, et al.
Published: (2025)
Tools Fail: Detecting Silent Errors in Faulty Tools
by: Sun, Jimin, et al.
Published: (2024)
by: Sun, Jimin, et al.
Published: (2024)
Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws
by: Sardana, Nikhil, et al.
Published: (2023)
by: Sardana, Nikhil, et al.
Published: (2023)
AE-LLM: Adaptive Efficiency Optimization for Large Language Models
by: Tanaka, Kaito, et al.
Published: (2026)
by: Tanaka, Kaito, et al.
Published: (2026)
SpreadsheetArena: Decomposing Preference in LLM Generation of Spreadsheet Workbooks
by: Kundurthy, Srivatsa, et al.
Published: (2026)
by: Kundurthy, Srivatsa, et al.
Published: (2026)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024)
by: Tyukin, Georgy
Published: (2024)
Towards Resource-Efficient LLMs: End-to-End Energy Accounting of Distillation Pipelines
by: Lambert, Katherine, et al.
Published: (2026)
by: Lambert, Katherine, et al.
Published: (2026)
Self-Regulation and Requesting Interventions
by: Min, So Yeon, et al.
Published: (2025)
by: Min, So Yeon, et al.
Published: (2025)
Sequence-Level Leakage Risk of Training Data in Large Language Models
by: Tiwari, Trishita, et al.
Published: (2024)
by: Tiwari, Trishita, et al.
Published: (2024)
Sparse Upcycling: Inference Inefficient Finetuning
by: Doubov, Sasha, et al.
Published: (2024)
by: Doubov, Sasha, et al.
Published: (2024)
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
by: Delavande, Julien, et al.
Published: (2025)
by: Delavande, Julien, et al.
Published: (2025)
Kinetics: Rethinking Test-Time Scaling Laws
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
by: Sadhukhan, Ranajoy, et al.
Published: (2025)
Dynamic Subset Tuning: Expanding the Operational Range of Parameter-Efficient Training for Large Language Models
by: Stahlberg, Felix, et al.
Published: (2024)
by: Stahlberg, Felix, et al.
Published: (2024)
Learning Model Successors
by: Chang, Yingshan, et al.
Published: (2025)
by: Chang, Yingshan, et al.
Published: (2025)
Strategic Polysemy in AI Discourse: A Philosophical Analysis of Language, Hype, and Power
by: LaCroix, Travis, et al.
Published: (2026)
by: LaCroix, Travis, et al.
Published: (2026)
Small Talk, Big Impact: The Energy Cost of Thanking AI
by: Delavande, Julien, et al.
Published: (2026)
by: Delavande, Julien, et al.
Published: (2026)
Optimization Strategies for Enhancing Resource Efficiency in Transformers & Large Language Models
by: Wallace, Tom, et al.
Published: (2025)
by: Wallace, Tom, et al.
Published: (2025)
CATS: Contextually-Aware Thresholding for Sparsity in Large Language Models
by: Lee, Donghyun, et al.
Published: (2024)
by: Lee, Donghyun, et al.
Published: (2024)
CORM: Cache Optimization with Recent Message for Large Language Model Inference
by: Dai, Jincheng, et al.
Published: (2024)
by: Dai, Jincheng, et al.
Published: (2024)
Optimizing Large Language Models with an Enhanced LoRA Fine-Tuning Algorithm for Efficiency and Robustness in NLP Tasks
by: Hu, Jiacheng, et al.
Published: (2024)
by: Hu, Jiacheng, et al.
Published: (2024)
Under the Influence: Quantifying Persuasion and Vigilance in Large Language Models
by: Robinson, Sasha, et al.
Published: (2026)
by: Robinson, Sasha, et al.
Published: (2026)
Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning
by: Sclar, Melanie, et al.
Published: (2024)
by: Sclar, Melanie, et al.
Published: (2024)
Adversarial Evasion Attack Efficiency against Large Language Models
by: Vitorino, João, et al.
Published: (2024)
by: Vitorino, João, et al.
Published: (2024)
Inference-Cost-Aware Dynamic Tree Construction for Efficient Inference in Large Language Models
by: Hong, Yinrong, et al.
Published: (2025)
by: Hong, Yinrong, et al.
Published: (2025)
TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
by: Li, Yizhi, et al.
Published: (2025)
by: Li, Yizhi, et al.
Published: (2025)
Task-Specific Efficiency Analysis: When Small Language Models Outperform Large Language Models
by: Cao, Jinghan, et al.
Published: (2026)
by: Cao, Jinghan, et al.
Published: (2026)
FlashDecoding++: Faster Large Language Model Inference on GPUs
by: Hong, Ke, et al.
Published: (2023)
by: Hong, Ke, et al.
Published: (2023)
Experimental Design for Active Transductive Inference in Large Language Models
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
by: Mukherjee, Subhojyoti, et al.
Published: (2024)
Toward Sustainable GenAI using Generation Directives for Carbon-Friendly Large Language Model Inference
by: Li, Baolin, et al.
Published: (2024)
by: Li, Baolin, et al.
Published: (2024)
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
SOTOPIA: Interactive Evaluation for Social Intelligence in Language Agents
by: Zhou, Xuhui, et al.
Published: (2023)
by: Zhou, Xuhui, et al.
Published: (2023)
Similar Items
-
Energy and Carbon Considerations of Fine-Tuning BERT
by: Wang, Xiaorong, et al.
Published: (2023) -
Gradient Localization Improves Lifelong Pretraining of Language Models
by: Fernandez, Jared, et al.
Published: (2024) -
Language Models Need Inductive Biases to Count Inductively
by: Chang, Yingshan, et al.
Published: (2024) -
Power Hungry Processing: Watts Driving the Cost of AI Deployment?
by: Luccioni, Alexandra Sasha, et al.
Published: (2023) -
Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training
by: Fernandez, Jared, et al.
Published: (2024)