Energy-Aware LLMs: A step towards sustainable AI for downstream applications
Fuente:
arXiv
Saved in:
| Main Authors: | Tran, Nguyen Phuc, Jaumard, Brigitte, Delgado, Oscar |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proactive Service Assurance in 5G and B5G Networks: A Closed-Loop Algorithm for End-to-End Network Slicing
by: Tran, Nguyen Phuc, et al.
Published: (2024)
by: Tran, Nguyen Phuc, et al.
Published: (2024)
LLM-Augmented Knowledge Base Construction For Root Cause Analysis
by: Tran, Nguyen Phuc, et al.
Published: (2026)
by: Tran, Nguyen Phuc, et al.
Published: (2026)
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026)
by: Jha, Saurav, et al.
Published: (2026)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025)
by: Israel, Daniel, et al.
Published: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)
by: Wang, Wenxiao, et al.
Published: (2024)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
by: Zhao, Qihao, et al.
Published: (2024)
by: Zhao, Qihao, et al.
Published: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
by: Chiang, Hung-Yueh, et al.
Published: (2025)
by: Chiang, Hung-Yueh, et al.
Published: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
by: Lin, Yujun, et al.
Published: (2024)
by: Lin, Yujun, et al.
Published: (2024)
Data Efficacy for Language Model Training
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
by: Stuhlmann, Linus, et al.
Published: (2025)
by: Stuhlmann, Linus, et al.
Published: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
by: Bhope, Rahul Atul, et al.
Published: (2025)
by: Bhope, Rahul Atul, et al.
Published: (2025)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
by: Zandieh, Amir, et al.
Published: (2024)
by: Zandieh, Amir, et al.
Published: (2024)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
by: Amujo, Oluyemi Enoch, et al.
Published: (2024)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
by: Gultekin, Sinan, et al.
Published: (2023)
by: Gultekin, Sinan, et al.
Published: (2023)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
by: Tyukin, Georgy
Published: (2024)
by: Tyukin, Georgy
Published: (2024)
Distillation Contrastive Decoding: Improving LLMs Reasoning with Contrastive Decoding and Distillation
by: Phan, Phuc, et al.
Published: (2024)
by: Phan, Phuc, et al.
Published: (2024)
QoS-Aware Dynamic CU Selection in O-RAN with Graph-Based Reinforcement Learning
by: Racedo, Sebastian, et al.
Published: (2025)
by: Racedo, Sebastian, et al.
Published: (2025)
EXAQ: Exponent Aware Quantization For LLMs Acceleration
by: Shkolnik, Moran, et al.
Published: (2024)
by: Shkolnik, Moran, et al.
Published: (2024)
MoEITS: A Green AI approach for simplifying MoE-LLMs
by: Balderas, Luis, et al.
Published: (2026)
by: Balderas, Luis, et al.
Published: (2026)
Energy per Successful Goal: Goal-Level Energy Accounting for Agentic AI Systems
by: Panigrahy, Deepak, et al.
Published: (2026)
by: Panigrahy, Deepak, et al.
Published: (2026)
Regression Language Models for Code
by: Akhauri, Yash, et al.
Published: (2025)
by: Akhauri, Yash, et al.
Published: (2025)
CDS4RAG: Cyclic Dual-Sequential Hyperparameter Optimization for RAG
by: Chen, Pengzhou, et al.
Published: (2026)
by: Chen, Pengzhou, et al.
Published: (2026)
LLMSYS-HPOBench: Hyperparameter Optimization Benchmark Suite for Real-World LLM Systems
by: Wu, Siyu, et al.
Published: (2026)
by: Wu, Siyu, et al.
Published: (2026)
ML KPI Prediction in 5G and B5G Networks
by: Tran, Nguyen Phuc, et al.
Published: (2024)
by: Tran, Nguyen Phuc, et al.
Published: (2024)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
by: Pan, Haolin, et al.
Published: (2026)
by: Pan, Haolin, et al.
Published: (2026)
HashEvict: A Pre-Attention KV Cache Eviction Strategy using Locality-Sensitive Hashing
by: Liu, Minghui, et al.
Published: (2024)
by: Liu, Minghui, et al.
Published: (2024)
Metrics and evaluations for computational and sustainable AI efficiency
by: Liu, Hongyuan, et al.
Published: (2025)
by: Liu, Hongyuan, et al.
Published: (2025)
The Illusion of Power Capping in LLM Decode: A Phase-Aware Energy Characterisation Across Attention Architectures
by: Ma, Bole, et al.
Published: (2026)
by: Ma, Bole, et al.
Published: (2026)
Enhancing Energy-Awareness in Deep Learning through Fine-Grained Energy Measurement
by: Rajput, Saurabhsingh, et al.
Published: (2023)
by: Rajput, Saurabhsingh, et al.
Published: (2023)
R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token Routing
by: Fu, Tianyu, et al.
Published: (2025)
by: Fu, Tianyu, et al.
Published: (2025)
LLMs for Analog Circuit Design Continuum (ACDC)
by: Esfandiari, Yasaman, et al.
Published: (2025)
by: Esfandiari, Yasaman, et al.
Published: (2025)
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
by: Yang, Shang, et al.
Published: (2025)
by: Yang, Shang, et al.
Published: (2025)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
by: Zhang, Qizheng, et al.
Published: (2025)
by: Zhang, Qizheng, et al.
Published: (2025)
Canvas: End-to-End Kernel Architecture Search in Neural Networks
by: Zhao, Chenggang, et al.
Published: (2023)
by: Zhao, Chenggang, et al.
Published: (2023)
PATCH: Learnable Tile-level Hybrid Sparsity for LLMs
by: Hourri, Younes, et al.
Published: (2025)
by: Hourri, Younes, et al.
Published: (2025)
CORI: CJKV Benchmark with Romanization Integration -- A step towards Cross-lingual Transfer Beyond Textual Scripts
by: Nguyen, Hoang H., et al.
Published: (2024)
by: Nguyen, Hoang H., et al.
Published: (2024)
An MLCommons Scientific Benchmarks Ontology
by: Hawks, Ben, et al.
Published: (2025)
by: Hawks, Ben, et al.
Published: (2025)
The Race to Efficiency: A New Perspective on AI Scaling Laws
by: Lu, Chien-Ping
Published: (2025)
by: Lu, Chien-Ping
Published: (2025)
Similar Items
-
Proactive Service Assurance in 5G and B5G Networks: A Closed-Loop Algorithm for End-to-End Network Slicing
by: Tran, Nguyen Phuc, et al.
Published: (2024) -
LLM-Augmented Knowledge Base Construction For Root Cause Analysis
by: Tran, Nguyen Phuc, et al.
Published: (2026) -
REAM: Merging Improves Pruning of Experts in LLMs
by: Jha, Saurav, et al.
Published: (2026) -
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
by: Israel, Daniel, et al.
Published: (2025) -
Model Compression and Efficient Inference for Large Language Models: A Survey
by: Wang, Wenxiao, et al.
Published: (2024)