Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Yi, Qian, Wenbo, Xie, Dedong, Hu, Ruiyan, Hu, Yigong, Kasikci, Baris |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025)
by: Zhu, Kan, et al.
Published: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
by: Zarkadas, Ioannis, et al.
Published: (2025)
by: Zarkadas, Ioannis, et al.
Published: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025)
by: Sun, Wenbo, et al.
Published: (2025)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
by: Oviedo, Felipe, et al.
Published: (2025)
by: Oviedo, Felipe, et al.
Published: (2025)
Federated Communication-Efficient Multi-Objective Optimization
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
Intelligent Task Offloading in VANETs: A Hybrid AI-Driven Approach for Low-Latency and Energy Efficiency
by: Qayyum, Tariq, et al.
Published: (2025)
by: Qayyum, Tariq, et al.
Published: (2025)
A Scalable Digital Twin Framework for Energy Optimization in Data Centers
by: Gonçalves, Raphael Hendrigo de Souza, et al.
Published: (2026)
by: Gonçalves, Raphael Hendrigo de Souza, et al.
Published: (2026)
Communication and Energy Efficient Federated Learning using Zero-Order Optimization Technique
by: Mhanna, Elissa, et al.
Published: (2024)
by: Mhanna, Elissa, et al.
Published: (2024)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
by: Yao, Yuhang, et al.
Published: (2024)
by: Yao, Yuhang, et al.
Published: (2024)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
by: Kamahori, Keisuke, et al.
Published: (2024)
by: Kamahori, Keisuke, et al.
Published: (2024)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
by: Lin, Chien-Yu, et al.
Published: (2025)
by: Lin, Chien-Yu, et al.
Published: (2025)
BGTplanner: Maximizing Training Accuracy for Differentially Private Federated Recommenders via Strategic Privacy Budget Allocation
by: Zhang, Xianzhi, et al.
Published: (2024)
by: Zhang, Xianzhi, et al.
Published: (2024)
Toward Cross-Layer Energy Optimizations in AI Systems
by: Chung, Jae-Won, et al.
Published: (2024)
by: Chung, Jae-Won, et al.
Published: (2024)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
by: Dai, Yinwei, et al.
Published: (2023)
by: Dai, Yinwei, et al.
Published: (2023)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
by: Tschand, Arya, et al.
Published: (2024)
by: Tschand, Arya, et al.
Published: (2024)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
by: Kamahori, Keisuke, et al.
Published: (2026)
by: Kamahori, Keisuke, et al.
Published: (2026)
MatchNAS: Optimizing Edge AI in Sparse-Label Data Contexts via Automating Deep Neural Network Porting for Mobile Deployment
by: Huang, Hongtao, et al.
Published: (2024)
by: Huang, Hongtao, et al.
Published: (2024)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
by: Padovani, Gabriele, et al.
Published: (2025)
by: Padovani, Gabriele, et al.
Published: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
by: Zhang, Zongshun, et al.
Published: (2025)
by: Zhang, Zongshun, et al.
Published: (2025)
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
by: Zhang, Yanqi, et al.
Published: (2024)
by: Zhang, Yanqi, et al.
Published: (2024)
Deep Reinforcement Learning for Optimizing Energy Consumption in Smart Grid Systems
by: Alsheikhi, Abeer, et al.
Published: (2026)
by: Alsheikhi, Abeer, et al.
Published: (2026)
Reducing Energy Bloat in Large Model Training
by: Chung, Jae-Won, et al.
Published: (2023)
by: Chung, Jae-Won, et al.
Published: (2023)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
by: Pan, Yi, et al.
Published: (2026)
by: Pan, Yi, et al.
Published: (2026)
AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT Systems
by: Jia, Chentao, et al.
Published: (2023)
by: Jia, Chentao, et al.
Published: (2023)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
by: Gu, Yile, et al.
Published: (2025)
by: Gu, Yile, et al.
Published: (2025)
Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities
by: Wei, Yunze, et al.
Published: (2024)
by: Wei, Yunze, et al.
Published: (2024)
High-Dimensional Sparse Data Low-rank Representation via Accelerated Asynchronous Parallel Stochastic Gradient Descent
by: Hu, Qicong, et al.
Published: (2024)
by: Hu, Qicong, et al.
Published: (2024)
Energy-Efficient Federated Learning for Edge Real-Time Vision via Joint Data, Computation, and Communication Design
by: Hou, Xiangwang, et al.
Published: (2025)
by: Hou, Xiangwang, et al.
Published: (2025)
Energy-Aware Decentralized Learning with Intermittent Model Training
by: Dhasade, Akash, et al.
Published: (2024)
by: Dhasade, Akash, et al.
Published: (2024)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
by: Su, Jianchang, et al.
Published: (2026)
by: Su, Jianchang, et al.
Published: (2026)
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
by: Wen, Wei, et al.
Published: (2024)
by: Wen, Wei, et al.
Published: (2024)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
by: Chung, Jae-Won, et al.
Published: (2026)
by: Chung, Jae-Won, et al.
Published: (2026)
FedZero: Leveraging Renewable Excess Energy in Federated Learning
by: Wiesner, Philipp, et al.
Published: (2023)
by: Wiesner, Philipp, et al.
Published: (2023)
Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients
by: Zhang, Zikai, et al.
Published: (2024)
by: Zhang, Zikai, et al.
Published: (2024)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
by: Tian, Jian, et al.
Published: (2025)
by: Tian, Jian, et al.
Published: (2025)
CaBaFL: Asynchronous Federated Learning via Hierarchical Cache and Feature Balance
by: Xia, Zeke, et al.
Published: (2024)
by: Xia, Zeke, et al.
Published: (2024)
FedAST: Federated Asynchronous Simultaneous Training
by: Askin, Baris, et al.
Published: (2024)
by: Askin, Baris, et al.
Published: (2024)
Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
by: Wu, Ruofan, et al.
Published: (2026)
by: Wu, Ruofan, et al.
Published: (2026)
Energy-Efficient Quantized Federated Learning for Resource-constrained IoT devices
by: Compaoré, Wilfrid Sougrinoma, et al.
Published: (2025)
by: Compaoré, Wilfrid Sougrinoma, et al.
Published: (2025)
Similar Items
-
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
by: Gu, Yile, et al.
Published: (2025) -
PolyServe: Efficient Multi-SLO Serving at Scale
by: Zhu, Kan, et al.
Published: (2025) -
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
by: Zarkadas, Ioannis, et al.
Published: (2025) -
Ilargi: a GPU Compatible Factorized ML Model Training Framework
by: Sun, Wenbo, et al.
Published: (2025) -
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
by: Oviedo, Felipe, et al.
Published: (2025)