PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mehboob, Talha, Guo, Luanzheng, Tallent, Nathan, Zink, Michael, Irwin, David |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EcoLearn: Optimizing the Carbon Footprint of Federated Learning
von: Mehboob, Talha, et al.
Veröffentlicht: (2023)
von: Mehboob, Talha, et al.
Veröffentlicht: (2023)
On The Reproducibility Limitations of RAG Systems
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025)
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
von: Sun, Minqiu, et al.
Veröffentlicht: (2026)
von: Sun, Minqiu, et al.
Veröffentlicht: (2026)
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
Datacenter Energy Optimized Power Profiles
von: Narayanaswamy, Sreedhar, et al.
Veröffentlicht: (2025)
von: Narayanaswamy, Sreedhar, et al.
Veröffentlicht: (2025)
Power Stabilization for AI Training Datacenters
von: Choukse, Esha, et al.
Veröffentlicht: (2025)
von: Choukse, Esha, et al.
Veröffentlicht: (2025)
DCGen 1.1 Technical Report: Generating Datacenter Configurations (including IT, Power, Cooling)
von: Gnibga, Wedan Emmanuel, et al.
Veröffentlicht: (2026)
von: Gnibga, Wedan Emmanuel, et al.
Veröffentlicht: (2026)
Designing Datacenter Power Delivery Hierarchies for the AI Era
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
von: Wilkins, Grant, et al.
Veröffentlicht: (2026)
Power- and Fragmentation-aware Online Scheduling for GPU Datacenters
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
von: Lettich, Francesco, et al.
Veröffentlicht: (2024)
EcoShift: Performance-Aware Power Management for Power-Constrained Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
von: Zheng, Zhong, et al.
Veröffentlicht: (2026)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
von: Xu, Guanbin, et al.
Veröffentlicht: (2026)
MassiveGNN: Efficient Training via Prefetching for Massively Connected Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2024)
Coordinated Power Management on Heterogeneous Systems
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
von: Zheng, Zhong, et al.
Veröffentlicht: (2025)
NOMAD: Generating Embeddings for Massive Distributed Graphs
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
von: Sarkar, Aishwarya, et al.
Veröffentlicht: (2026)
CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter Training
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
von: Chen, Tiancheng, et al.
Veröffentlicht: (2025)
On the Convergence of Malleability and the HPC PowerStack: Exploiting Dynamism in Over-Provisioned and Power-Constrained HPC Systems
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
von: Arima, Eishi, et al.
Veröffentlicht: (2024)
Modeling the Impact of Fiber Latency on Compute-Communication Overlap in Geo-Distributed Multi-Datacenter AI Training
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
von: Papavasileiou, Ioannis, et al.
Veröffentlicht: (2026)
Distribution and Management of Datacenter Load Decoupling
von: Lin, Liuzixuan, et al.
Veröffentlicht: (2025)
von: Lin, Liuzixuan, et al.
Veröffentlicht: (2025)
Serving Compound Inference Systems on Datacenter GPUs
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
von: Devata, Sriram, et al.
Veröffentlicht: (2026)
Distributed Order Recording Techniques for Efficient Record-and-Replay of Multi-threaded Programs
von: Fu, Xiang, et al.
Veröffentlicht: (2026)
von: Fu, Xiang, et al.
Veröffentlicht: (2026)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
von: Jang, Insu, et al.
Veröffentlicht: (2026)
von: Jang, Insu, et al.
Veröffentlicht: (2026)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhong, Yuchen, et al.
Veröffentlicht: (2024)
Overcoming Memory Constraints in Quantum Circuit Simulation with a High-Fidelity Compression Framework
von: Zhang, Boyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Boyuan, et al.
Veröffentlicht: (2024)
FEPLB: Exploiting Copy Engines for Nearly Free MoE Load Balancing in Distributed Training
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
von: Qi, Shuyao, et al.
Veröffentlicht: (2026)
CARAT: Client-Side Adaptive RPC and Cache Co-Tuning for Parallel File Systems
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026)
DFPL: Decentralized Federated Prototype Learning Across Heterogeneous Data Distributions
von: Zhang, Hongliang, et al.
Veröffentlicht: (2025)
von: Zhang, Hongliang, et al.
Veröffentlicht: (2025)
Scrutinizing Variables for Checkpoint Using Automatic Differentiation
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
Speedup of Distributed Algorithms for Power Graphs in the CONGEST Model
von: Barenboim, Leonid, et al.
Veröffentlicht: (2023)
von: Barenboim, Leonid, et al.
Veröffentlicht: (2023)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
von: Lv, Cunchi, et al.
Veröffentlicht: (2025)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
von: Zhang, WenZheng, et al.
Veröffentlicht: (2024)
Generic and ML Workloads in an HPC Datacenter: Node Energy, Job Failures, and Node-Job Analysis
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Chu, Xiaoyu, et al.
Veröffentlicht: (2024)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
von: Svedas, Jonas, et al.
Veröffentlicht: (2026)
Tetris: Efficient Intra-Datacenter Calls Packing for Large Conferencing Services
von: Gandhi, Rohan, et al.
Veröffentlicht: (2025)
von: Gandhi, Rohan, et al.
Veröffentlicht: (2025)
To Offload or Not To Offload: Model-driven Comparison of Edge-native and On-device Processing In the Era of Accelerators
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
von: Ng, Nathan, et al.
Veröffentlicht: (2025)
Capsule: Efficient Player Isolation for Datacenters
von: Du, Zhouheng, et al.
Veröffentlicht: (2025)
von: Du, Zhouheng, et al.
Veröffentlicht: (2025)
Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed Clusters
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
von: Strati, Foteini, et al.
Veröffentlicht: (2025)
Designing Dense Satellite Clusters for Distributed Space-based Datacenters
von: Pénot, Jules, et al.
Veröffentlicht: (2026)
von: Pénot, Jules, et al.
Veröffentlicht: (2026)
The Ghost in the Datacenter: Link Flapping, Topology Knowledge Failures, and the FITO Category Mistake
von: Borrill, Paul
Veröffentlicht: (2026)
von: Borrill, Paul
Veröffentlicht: (2026)
Understanding Power Consumption Metric on Heterogeneous Memory Systems
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
von: Proaño, Andrès Rubio, et al.
Veröffentlicht: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
von: Adikari, Tharindu, et al.
Veröffentlicht: (2024)
von: Adikari, Tharindu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EcoLearn: Optimizing the Carbon Footprint of Federated Learning
von: Mehboob, Talha, et al.
Veröffentlicht: (2023) -
On The Reproducibility Limitations of RAG Systems
von: Wang, Baiqiang, et al.
Veröffentlicht: (2025) -
LLMTailor: A Layer-wise Tailoring Tool for Efficient Checkpointing of Large Language Models
von: Sun, Minqiu, et al.
Veröffentlicht: (2026) -
QoSFlow: Ensuring Service Quality of Distributed Workflows Using Interpretable Sensitivity Models
von: Rashid, Md Hasanur, et al.
Veröffentlicht: (2026) -
Datacenter Energy Optimized Power Profiles
von: Narayanaswamy, Sreedhar, et al.
Veröffentlicht: (2025)