Where Do the Joules Go? Diagnosing Inference Energy Consumption
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chung, Jae-Won, Wu, Ruofan, Ma, Jeff J., Chowdhury, Mosharaf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
von: Wu, Ruofan, et al.
Veröffentlicht: (2026)
von: Wu, Ruofan, et al.
Veröffentlicht: (2026)
Toward Cross-Layer Energy Optimizations in AI Systems
von: Chung, Jae-Won, et al.
Veröffentlicht: (2024)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2024)
Reducing Energy Bloat in Large Model Training
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023)
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
von: Ma, Jeff J., et al.
Veröffentlicht: (2025)
von: Ma, Jeff J., et al.
Veröffentlicht: (2025)
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
von: Liu, Jiachen, et al.
Veröffentlicht: (2024)
TetriServe: Efficient DiT Serving for Heterogeneous Image Generation
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
von: Lu, Runyu, et al.
Veröffentlicht: (2025)
OpenG2G: A Simulation Platform for AI Datacenter-Grid Runtime Coordination
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026)
Venn: Resource Management for Collaborative Learning Jobs
von: Liu, Jiachen, et al.
Veröffentlicht: (2023)
von: Liu, Jiachen, et al.
Veröffentlicht: (2023)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
von: Jang, Insu, et al.
Veröffentlicht: (2026)
von: Jang, Insu, et al.
Veröffentlicht: (2026)
FedTrans: Efficient Federated Learning via Multi-Model Transformation
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
von: Zhu, Yuxuan, et al.
Veröffentlicht: (2024)
SHARe-KAN: Post-Training Vector Quantization for Cache-Resident KAN Inference
von: Smith, Jeff
Veröffentlicht: (2025)
von: Smith, Jeff
Veröffentlicht: (2025)
Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
von: Zhang, Haolin, et al.
Veröffentlicht: (2025)
KAIROS: Stateful, Context-Aware Power-Efficient Agentic Inference Serving
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
von: Yuan, Yichao, et al.
Veröffentlicht: (2026)
Exploring the Boundaries of On-Device Inference: When Tiny Falls Short, Go Hierarchical
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
von: Behera, Adarsh Prasad, et al.
Veröffentlicht: (2024)
BEFL: Balancing Energy Consumption in Federated Learning for Mobile Edge IoT
von: Ju, Zehao, et al.
Veröffentlicht: (2024)
von: Ju, Zehao, et al.
Veröffentlicht: (2024)
Arctic Inference with Shift Parallelism: Fast and Efficient Open Source Inference System for Enterprise AI
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
von: Rajbhandari, Samyam, et al.
Veröffentlicht: (2025)
Optimization of Energy Consumption Forecasting in Puno using Parallel Computing and ARIMA Models: An Innovative Approach to Big Data Processing
von: Vilca-Tinta, Cliver W., et al.
Veröffentlicht: (2024)
von: Vilca-Tinta, Cliver W., et al.
Veröffentlicht: (2024)
Going Forward-Forward in Distributed Deep Learning
von: Aktemur, Ege, et al.
Veröffentlicht: (2024)
von: Aktemur, Ege, et al.
Veröffentlicht: (2024)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
von: Oviedo, Felipe, et al.
Veröffentlicht: (2025)
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
von: Gupta, Vima, et al.
Veröffentlicht: (2024)
Where is the Testbed for my Federated Learning Research?
von: Božič, Janez, et al.
Veröffentlicht: (2024)
von: Božič, Janez, et al.
Veröffentlicht: (2024)
Adaptive Resolution Inference (ARI): Energy-Efficient Machine Learning for Internet of Things
von: Wang, Ziheng, et al.
Veröffentlicht: (2024)
von: Wang, Ziheng, et al.
Veröffentlicht: (2024)
Efficient Distributed MLLM Training with Cornstarch
von: Jang, Insu, et al.
Veröffentlicht: (2025)
von: Jang, Insu, et al.
Veröffentlicht: (2025)
Go With The Flow: Churn-Tolerant Decentralized Training of Large Language Models
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
von: Blagoev, Nikolay, et al.
Veröffentlicht: (2025)
LoGoFair: Post-Processing for Local and Global Fairness in Federated Learning
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Accelerating Large Language Model Training with 4D Parallelism and Memory Consumption Estimator
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
von: Fujii, Kazuki, et al.
Veröffentlicht: (2024)
Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless Computing
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
von: Liu, Mengfan, et al.
Veröffentlicht: (2025)
Designing Large Foundation Models for Efficient Training and Inference: A Survey
von: Liu, Dong, et al.
Veröffentlicht: (2024)
von: Liu, Dong, et al.
Veröffentlicht: (2024)
CascadeServe: Unlocking Model Cascades for Inference Serving
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
von: Kossmann, Ferdi, et al.
Veröffentlicht: (2024)
Collaborative Speculative Inference for Efficient LLM Inference Serving
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
von: Gao, Luyao, et al.
Veröffentlicht: (2025)
A Survey on Collaborative DNN Inference for Edge Intelligence
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
von: Ren, Weiqing, et al.
Veröffentlicht: (2022)
Fast Distributed Inference Serving for Large Language Models
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
von: Wu, Bingyang, et al.
Veröffentlicht: (2023)
Floe: Federated Specialization for Real-Time LLM-SLM Inference
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
von: Tian, Chunlin, et al.
Veröffentlicht: (2026)
Quantifying the Energy Consumption and Carbon Emissions of LLM Inference via Simulations
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
von: Özcan, Miray, et al.
Veröffentlicht: (2025)
Deep Reinforcement Learning for Optimizing Energy Consumption in Smart Grid Systems
von: Alsheikhi, Abeer, et al.
Veröffentlicht: (2026)
von: Alsheikhi, Abeer, et al.
Veröffentlicht: (2026)
cuConv: A CUDA Implementation of Convolution for CNN Inference
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
von: Jordà, Marc, et al.
Veröffentlicht: (2021)
Compare Where It Matters: Using Layer-Wise Regularization To Improve Federated Learning on Heterogeneous Data
von: Son, Ha Min, et al.
Veröffentlicht: (2021)
von: Son, Ha Min, et al.
Veröffentlicht: (2021)
Salted Inference: Enhancing Privacy while Maintaining Efficiency of Split Inference in Mobile Computing
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
von: Malekzadeh, Mohammad, et al.
Veröffentlicht: (2023)
MoE-Compression: How the Compression Error of Experts Affects the Inference Accuracy of MoE Model?
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
von: Ma, Songkai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
von: Wu, Ruofan, et al.
Veröffentlicht: (2026) -
Toward Cross-Layer Energy Optimizations in AI Systems
von: Chung, Jae-Won, et al.
Veröffentlicht: (2024) -
Reducing Energy Bloat in Large Model Training
von: Chung, Jae-Won, et al.
Veröffentlicht: (2023) -
Cornserve: A Distributed Serving System for Any-to-Any Multimodal Models
von: Chung, Jae-Won, et al.
Veröffentlicht: (2026) -
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
von: Ma, Jeff J., et al.
Veröffentlicht: (2025)