Magneton: Optimizing Energy Efficiency of ML Systems via Differential Energy Debugging
Fuente:
arXiv
Guardado en:
| Autores principales: | Pan, Yi, Qian, Wenbo, Xie, Dedong, Hu, Ruiyan, Hu, Yigong, Kasikci, Baris |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
por: Gu, Yile, et al.
Publicado: (2025)
por: Gu, Yile, et al.
Publicado: (2025)
PolyServe: Efficient Multi-SLO Serving at Scale
por: Zhu, Kan, et al.
Publicado: (2025)
por: Zhu, Kan, et al.
Publicado: (2025)
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
por: Zarkadas, Ioannis, et al.
Publicado: (2025)
por: Zarkadas, Ioannis, et al.
Publicado: (2025)
Ilargi: a GPU Compatible Factorized ML Model Training Framework
por: Sun, Wenbo, et al.
Publicado: (2025)
por: Sun, Wenbo, et al.
Publicado: (2025)
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
por: Oviedo, Felipe, et al.
Publicado: (2025)
por: Oviedo, Felipe, et al.
Publicado: (2025)
Federated Communication-Efficient Multi-Objective Optimization
por: Askin, Baris, et al.
Publicado: (2024)
por: Askin, Baris, et al.
Publicado: (2024)
Intelligent Task Offloading in VANETs: A Hybrid AI-Driven Approach for Low-Latency and Energy Efficiency
por: Qayyum, Tariq, et al.
Publicado: (2025)
por: Qayyum, Tariq, et al.
Publicado: (2025)
A Scalable Digital Twin Framework for Energy Optimization in Data Centers
por: Gonçalves, Raphael Hendrigo de Souza, et al.
Publicado: (2026)
por: Gonçalves, Raphael Hendrigo de Souza, et al.
Publicado: (2026)
Communication and Energy Efficient Federated Learning using Zero-Order Optimization Technique
por: Mhanna, Elissa, et al.
Publicado: (2024)
por: Mhanna, Elissa, et al.
Publicado: (2024)
ScaleLLM: A Resource-Frugal LLM Serving Framework by Optimizing End-to-End Efficiency
por: Yao, Yuhang, et al.
Publicado: (2024)
por: Yao, Yuhang, et al.
Publicado: (2024)
Fiddler: CPU-GPU Orchestration for Fast Inference of Mixture-of-Experts Models
por: Kamahori, Keisuke, et al.
Publicado: (2024)
por: Kamahori, Keisuke, et al.
Publicado: (2024)
TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
por: Lin, Chien-Yu, et al.
Publicado: (2025)
por: Lin, Chien-Yu, et al.
Publicado: (2025)
BGTplanner: Maximizing Training Accuracy for Differentially Private Federated Recommenders via Strategic Privacy Budget Allocation
por: Zhang, Xianzhi, et al.
Publicado: (2024)
por: Zhang, Xianzhi, et al.
Publicado: (2024)
Toward Cross-Layer Energy Optimizations in AI Systems
por: Chung, Jae-Won, et al.
Publicado: (2024)
por: Chung, Jae-Won, et al.
Publicado: (2024)
Apparate: Rethinking Early Exits to Tame Latency-Throughput Tensions in ML Serving
por: Dai, Yinwei, et al.
Publicado: (2023)
por: Dai, Yinwei, et al.
Publicado: (2023)
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
por: Tschand, Arya, et al.
Publicado: (2024)
por: Tschand, Arya, et al.
Publicado: (2024)
VibeServe: Can AI Agents Build Bespoke LLM Serving Systems?
por: Kamahori, Keisuke, et al.
Publicado: (2026)
por: Kamahori, Keisuke, et al.
Publicado: (2026)
MatchNAS: Optimizing Edge AI in Sparse-Label Data Contexts via Automating Deep Neural Network Porting for Mobile Deployment
por: Huang, Hongtao, et al.
Publicado: (2024)
por: Huang, Hongtao, et al.
Publicado: (2024)
yProv4ML: Effortless Provenance Tracking for Machine Learning Systems
por: Padovani, Gabriele, et al.
Publicado: (2025)
por: Padovani, Gabriele, et al.
Publicado: (2025)
SERFLOW: A Cross-Service Cost Optimization Framework for SLO-Aware Dynamic ML Inference
por: Zhang, Zongshun, et al.
Publicado: (2025)
por: Zhang, Zongshun, et al.
Publicado: (2025)
DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV Compaction
por: Zhang, Yanqi, et al.
Publicado: (2024)
por: Zhang, Yanqi, et al.
Publicado: (2024)
Deep Reinforcement Learning for Optimizing Energy Consumption in Smart Grid Systems
por: Alsheikhi, Abeer, et al.
Publicado: (2026)
por: Alsheikhi, Abeer, et al.
Publicado: (2026)
Reducing Energy Bloat in Large Model Training
por: Chung, Jae-Won, et al.
Publicado: (2023)
por: Chung, Jae-Won, et al.
Publicado: (2023)
DynaFlow: Transparent and Flexible Intra-Device Parallelism via Programmable Operator Scheduling
por: Pan, Yi, et al.
Publicado: (2026)
por: Pan, Yi, et al.
Publicado: (2026)
AdaptiveFL: Adaptive Heterogeneous Federated Learning for Resource-Constrained AIoT Systems
por: Jia, Chentao, et al.
Publicado: (2023)
por: Jia, Chentao, et al.
Publicado: (2023)
ConsumerBench: Benchmarking Generative AI Applications on End-User Devices
por: Gu, Yile, et al.
Publicado: (2025)
por: Gu, Yile, et al.
Publicado: (2025)
Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities
por: Wei, Yunze, et al.
Publicado: (2024)
por: Wei, Yunze, et al.
Publicado: (2024)
High-Dimensional Sparse Data Low-rank Representation via Accelerated Asynchronous Parallel Stochastic Gradient Descent
por: Hu, Qicong, et al.
Publicado: (2024)
por: Hu, Qicong, et al.
Publicado: (2024)
Energy-Efficient Federated Learning for Edge Real-Time Vision via Joint Data, Computation, and Communication Design
por: Hou, Xiangwang, et al.
Publicado: (2025)
por: Hou, Xiangwang, et al.
Publicado: (2025)
Energy-Aware Decentralized Learning with Intermittent Model Training
por: Dhasade, Akash, et al.
Publicado: (2024)
por: Dhasade, Akash, et al.
Publicado: (2024)
SAIR: Cost-Efficient Multi-Stage ML Pipeline Autoscaling via In-Context Reinforcement Learning
por: Su, Jianchang, et al.
Publicado: (2026)
por: Su, Jianchang, et al.
Publicado: (2026)
CubicML: Automated ML for Large ML Systems Co-design with ML Prediction of Performance
por: Wen, Wei, et al.
Publicado: (2024)
por: Wen, Wei, et al.
Publicado: (2024)
Where Do the Joules Go? Diagnosing Inference Energy Consumption
por: Chung, Jae-Won, et al.
Publicado: (2026)
por: Chung, Jae-Won, et al.
Publicado: (2026)
FedZero: Leveraging Renewable Excess Energy in Federated Learning
por: Wiesner, Philipp, et al.
Publicado: (2023)
por: Wiesner, Philipp, et al.
Publicado: (2023)
Fed-pilot: Optimizing LoRA Allocation for Efficient Federated Fine-Tuning with Heterogeneous Clients
por: Zhang, Zikai, et al.
Publicado: (2024)
por: Zhang, Zikai, et al.
Publicado: (2024)
Staggered Batch Scheduling: Co-optimizing Time-to-First-Token and Throughput for High-Efficiency LLM Inference
por: Tian, Jian, et al.
Publicado: (2025)
por: Tian, Jian, et al.
Publicado: (2025)
CaBaFL: Asynchronous Federated Learning via Hierarchical Cache and Feature Balance
por: Xia, Zeke, et al.
Publicado: (2024)
por: Xia, Zeke, et al.
Publicado: (2024)
FedAST: Federated Asynchronous Simultaneous Training
por: Askin, Baris, et al.
Publicado: (2024)
por: Askin, Baris, et al.
Publicado: (2024)
Kareus: Joint Reduction of Dynamic and Static Energy in Large Model Training
por: Wu, Ruofan, et al.
Publicado: (2026)
por: Wu, Ruofan, et al.
Publicado: (2026)
Energy-Efficient Quantized Federated Learning for Resource-constrained IoT devices
por: Compaoré, Wilfrid Sougrinoma, et al.
Publicado: (2025)
por: Compaoré, Wilfrid Sougrinoma, et al.
Publicado: (2025)
Ejemplares similares
-
Argos: Agentic Time-Series Anomaly Detection with Autonomous Rule Generation via Large Language Models
por: Gu, Yile, et al.
Publicado: (2025) -
PolyServe: Efficient Multi-SLO Serving at Scale
por: Zhu, Kan, et al.
Publicado: (2025) -
Fake Runs, Real Fixes -- Analyzing xPU Performance Through Simulation
por: Zarkadas, Ioannis, et al.
Publicado: (2025) -
Ilargi: a GPU Compatible Factorized ML Model Training Framework
por: Sun, Wenbo, et al.
Publicado: (2025) -
Energy Use of AI Inference: Efficiency Pathways and Test-Time Compute
por: Oviedo, Felipe, et al.
Publicado: (2025)