Characterizing the Efficiency of Distributed Training: A Power, Performance, and Thermal Perspective
Fuente:
arXiv
Saved in:
| Main Authors: | Go, Seokjin, Park, Joongun, More, Spandan, Wu, Hanjiang, Wang, Irene, Jezghani, Aaron, Krishna, Tushar, Mahajan, Divya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
by: Lee, Seonho, et al.
Published: (2025)
by: Lee, Seonho, et al.
Published: (2025)
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
by: Go, Seokjin, et al.
Published: (2025)
by: Go, Seokjin, et al.
Published: (2025)
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
by: Man, Changhai, et al.
Published: (2025)
by: Man, Changhai, et al.
Published: (2025)
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
by: Go, Seokjin, et al.
Published: (2026)
by: Go, Seokjin, et al.
Published: (2026)
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
by: Wu, Hanjiang, et al.
Published: (2024)
by: Wu, Hanjiang, et al.
Published: (2024)
Characterizing CPU-Induced Slowdowns in Multi-GPU LLM Inference
by: Chung, Euijun, et al.
Published: (2026)
by: Chung, Euijun, et al.
Published: (2026)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Workload-Aware Hardware Accelerator Mining for Distributed Deep Learning Training
by: Adnan, Muhammad, et al.
Published: (2024)
by: Adnan, Muhammad, et al.
Published: (2024)
EarthSight: A Distributed Framework for Low-Latency Satellite Intelligence
by: Erol, Ansel Kaplan, et al.
Published: (2025)
by: Erol, Ansel Kaplan, et al.
Published: (2025)
LIBRA: Enabling Workload-aware Multi-dimensional Network Topology Optimization for Distributed Training of Large AI Models
by: Won, William, et al.
Published: (2021)
by: Won, William, et al.
Published: (2021)
Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
by: Svedas, Jonas, et al.
Published: (2026)
by: Svedas, Jonas, et al.
Published: (2026)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
by: Yoo, Jinsun, et al.
Published: (2026)
by: Yoo, Jinsun, et al.
Published: (2026)
MLCommons Chakra: Advancing Performance Benchmarking and Co-design using Standardized Execution Traces
by: Sridharan, Srinivas, et al.
Published: (2026)
by: Sridharan, Srinivas, et al.
Published: (2026)
PowerTrip: Exploiting Federated Heterogeneous Datacenter Power for Distributed ML Training
by: Mehboob, Talha, et al.
Published: (2025)
by: Mehboob, Talha, et al.
Published: (2025)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
by: K., Prashanthi S., et al.
Published: (2023)
by: K., Prashanthi S., et al.
Published: (2023)
Lagom: Unleashing the Power of Communication and Computation Overlapping for Distributed LLM Training
by: Xu, Guanbin, et al.
Published: (2026)
by: Xu, Guanbin, et al.
Published: (2026)
Integrated Hardware Architecture and Device Placement Search
by: Wang, Irene, et al.
Published: (2024)
by: Wang, Irene, et al.
Published: (2024)
COSMIC: Enabling Full-Stack Co-Design and Optimization of Distributed Machine Learning Systems
by: Raju, Aditi, et al.
Published: (2025)
by: Raju, Aditi, et al.
Published: (2025)
Comparative Analysis of Lightweight Kubernetes Distributions for Edge Computing: Performance and Resource Efficiency
by: Yakubov, Diyaz, et al.
Published: (2025)
by: Yakubov, Diyaz, et al.
Published: (2025)
On the Performance and Memory Footprint of Distributed Training: An Empirical Study on Transformers
by: Lu, Zhengxian, et al.
Published: (2024)
by: Lu, Zhengxian, et al.
Published: (2024)
Characterizing the Performance of Accelerated Jetson Edge Devices for Training Deep Learning Models
by: K., Prashanthi S., et al.
Published: (2025)
by: K., Prashanthi S., et al.
Published: (2025)
HopGNN: Boosting Distributed GNN Training Efficiency via Feature-Centric Model Migration
by: Chen, Weijian, et al.
Published: (2024)
by: Chen, Weijian, et al.
Published: (2024)
TACOS: Topology-Aware Collective Algorithm Synthesizer for Distributed Machine Learning
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
A Study on the Performance of Distributed Training of Data-driven CFD Simulations
by: Iserte, Sergio, et al.
Published: (2026)
by: Iserte, Sergio, et al.
Published: (2026)
Towards Cloud Efficiency with Large-scale Workload Characterization
by: Parayil, Anjaly, et al.
Published: (2024)
by: Parayil, Anjaly, et al.
Published: (2024)
HARP: A Taxonomy for Heterogeneous and Hierarchical Processors for Mixed-reuse Workloads
by: Garg, Raveesh, et al.
Published: (2025)
by: Garg, Raveesh, et al.
Published: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
Towards a Standardized Representation for Deep Learning Collective Algorithms
by: Yoo, Jinsun, et al.
Published: (2024)
by: Yoo, Jinsun, et al.
Published: (2024)
ASTRA-sim2.0: Modeling Hierarchical Networks and Disaggregated Systems for Large-model Training at Scale
by: Won, William, et al.
Published: (2023)
by: Won, William, et al.
Published: (2023)
Equinox: Decentralized Scheduling for Hardware-Aware Orbital Intelligence
by: Erol, Ansel Kaplan, et al.
Published: (2026)
by: Erol, Ansel Kaplan, et al.
Published: (2026)
Topological Characterization of Consensus in Distributed Systems
by: Nowak, Thomas, et al.
Published: (2019)
by: Nowak, Thomas, et al.
Published: (2019)
Privacy-Preserving Federated Learning: Integrating Zero-Knowledge Proofs in Scalable Distributed Architectures
by: Gupta, Divya
Published: (2026)
by: Gupta, Divya
Published: (2026)
Portability Efficiency Approach for Calculating Performance Portability
by: Marowka, Ami
Published: (2024)
by: Marowka, Ami
Published: (2024)
Zorse: Optimizing LLM Training Efficiency on Heterogeneous GPU Clusters
by: Guo, Runsheng Benson, et al.
Published: (2025)
by: Guo, Runsheng Benson, et al.
Published: (2025)
Exploring the Frontiers of Energy Efficiency using Power Management at System Scale
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
by: Karimi, Ahmad Maroof, et al.
Published: (2024)
Cloudless-Training: A Framework to Improve Efficiency of Geo-Distributed ML Training
by: Tan, Wenting, et al.
Published: (2023)
by: Tan, Wenting, et al.
Published: (2023)
Efficient Distributed MLLM Training with Cornstarch
by: Jang, Insu, et al.
Published: (2025)
by: Jang, Insu, et al.
Published: (2025)
Performance Characterization of Distributed Deep Learning Strategies: A Quantitative Evaluation of DDP, FSDP, and Parameter Server Architectures on GPU Clusters
by: Ovi, Md Sultanul Islam
Published: (2025)
by: Ovi, Md Sultanul Islam
Published: (2025)
The Power of Abstract MAC Layer: A Fault-tolerance Perspective
by: Zhang, Qinzi, et al.
Published: (2024)
by: Zhang, Qinzi, et al.
Published: (2024)
Similar Items
-
Characterizing Compute-Communication Overlap in GPU-Accelerated Distributed Deep Learning: Performance and Power Implications
by: Lee, Seonho, et al.
Published: (2025) -
MoETuner: Optimized Mixture of Expert Serving with Balanced Expert Placement and Token Routing
by: Go, Seokjin, et al.
Published: (2025) -
STAGE: A Symbolic Tensor grAph GEnerator for distributed AI system co-design
by: Man, Changhai, et al.
Published: (2025) -
ViBE: Co-Optimizing Workload Skew and Hardware Variability for MoE Serving
by: Go, Seokjin, et al.
Published: (2026) -
Enhancing Scalability and Performance in Influence Maximization with Optimized Parallel Processing
by: Wu, Hanjiang, et al.
Published: (2024)