Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
Fuente:
arXiv
Salvato in:
| Autori principali: | Yu, Enda, Dong, Dezun, Liao, Xiangke |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Nezha: A Key-Value Separated Distributed Store with Optimized Raft Integration
di: Wang, Yangyang, et al.
Pubblicazione: (2026)
di: Wang, Yangyang, et al.
Pubblicazione: (2026)
UNR: Unified Notifiable RMA Library for HPC
di: Feng, Guangnan, et al.
Pubblicazione: (2024)
di: Feng, Guangnan, et al.
Pubblicazione: (2024)
Demystifying ARM SME to Optimize General Matrix Multiplications
di: Deng, Chencheng, et al.
Pubblicazione: (2025)
di: Deng, Chencheng, et al.
Pubblicazione: (2025)
Distributed Set-membership Filtering Frameworks For Multi-agent Systems With Absolute and Relative Measurements
di: Ding, Yu, et al.
Pubblicazione: (2023)
di: Ding, Yu, et al.
Pubblicazione: (2023)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
di: Zhang, WenZheng, et al.
Pubblicazione: (2024)
di: Zhang, WenZheng, et al.
Pubblicazione: (2024)
A Survey of End-to-End Modeling for Distributed DNN Training: Workloads, Simulators, and TCO
di: Svedas, Jonas, et al.
Pubblicazione: (2025)
di: Svedas, Jonas, et al.
Pubblicazione: (2025)
HYDRA: Breaking the Global Ordering Barrier in Multi-BFT Consensus
di: Lyu, Hanzheng, et al.
Pubblicazione: (2025)
di: Lyu, Hanzheng, et al.
Pubblicazione: (2025)
Online Optimization of DNN Inference Network Utility in Collaborative Edge Computing
di: Li, Rui, et al.
Pubblicazione: (2024)
di: Li, Rui, et al.
Pubblicazione: (2024)
Will LLMs Scaling Hit the Wall? Breaking Barriers via Distributed Resources on Massive Edge Devices
di: Shen, Tao, et al.
Pubblicazione: (2025)
di: Shen, Tao, et al.
Pubblicazione: (2025)
Training DNN Models over Heterogeneous Clusters with Optimal Performance
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
di: Nie, Chengyi, et al.
Pubblicazione: (2024)
Parcae: Proactive, Liveput-Optimized DNN Training on Preemptible Instances
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
di: Duan, Jiangfei, et al.
Pubblicazione: (2024)
Performance Characterization of Containerized DNN Training and Inference on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
di: K., Prashanthi S., et al.
Pubblicazione: (2023)
Fulcrum: Optimizing Concurrent DNN Training and Inferencing on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
SWIFT: Expedited Failure Recovery for Large-scale DNN Training
di: Zhong, Yuchen, et al.
Pubblicazione: (2023)
di: Zhong, Yuchen, et al.
Pubblicazione: (2023)
A Flexible Programmable Pipeline Parallelism Framework for Efficient DNN Training
di: Jiang, Lijuan, et al.
Pubblicazione: (2025)
di: Jiang, Lijuan, et al.
Pubblicazione: (2025)
HAP: SPMD DNN Training on Heterogeneous GPU Clusters with Automated Program Synthesis
di: Zhang, Shiwei, et al.
Pubblicazione: (2024)
di: Zhang, Shiwei, et al.
Pubblicazione: (2024)
EdgeServing: Deadline-Aware Multi-DNN Serving at the Edge
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
di: Cao, Jiahe, et al.
Pubblicazione: (2026)
RailS: Load Balancing for All-to-All Communication in Distributed Mixture-of-Experts Training
di: Xu, Heng, et al.
Pubblicazione: (2025)
di: Xu, Heng, et al.
Pubblicazione: (2025)
EaCO: Resource Sharing Dynamics and Its Impact on Energy Efficiency for DNN Training
di: Haghshenas, Kawsar, et al.
Pubblicazione: (2024)
di: Haghshenas, Kawsar, et al.
Pubblicazione: (2024)
Evaluating Multi-Instance DNN Inferencing on Multiple Accelerators of an Edge Device
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
di: Tayal, Mumuksh, et al.
Pubblicazione: (2025)
Memory Efficient and Staleness Free Pipeline Parallel DNN Training Framework with Improved Convergence Speed
di: Dutta, Ankita, et al.
Pubblicazione: (2025)
di: Dutta, Ankita, et al.
Pubblicazione: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
di: Sun, Mo, et al.
Pubblicazione: (2024)
di: Sun, Mo, et al.
Pubblicazione: (2024)
AdaBridge: Dynamic Data and Computation Reuse for Efficient Multi-task DNN Co-evolution in Edge Systems
di: Wang, Lehao, et al.
Pubblicazione: (2024)
di: Wang, Lehao, et al.
Pubblicazione: (2024)
Towards a Flexible and High-Fidelity Approach to Distributed DNN Training Emulation
di: Liu, Banruo, et al.
Pubblicazione: (2024)
di: Liu, Banruo, et al.
Pubblicazione: (2024)
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching
di: Guo, Cong, et al.
Pubblicazione: (2024)
di: Guo, Cong, et al.
Pubblicazione: (2024)
TiMePReSt: Time and Memory Efficient Pipeline Parallel DNN Training with Removed Staleness
di: Dutta, Ankita, et al.
Pubblicazione: (2024)
di: Dutta, Ankita, et al.
Pubblicazione: (2024)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
di: Wang, Kewei, et al.
Pubblicazione: (2025)
di: Wang, Kewei, et al.
Pubblicazione: (2025)
IsoSched: Preemptive Tile Cascaded Scheduling of Multi-DNN via Subgraph Isomorphism
di: Zhao, Boran, et al.
Pubblicazione: (2025)
di: Zhao, Boran, et al.
Pubblicazione: (2025)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
di: Zhong, Yuchen, et al.
Pubblicazione: (2024)
di: Zhong, Yuchen, et al.
Pubblicazione: (2024)
AdaOper: Energy-efficient and Responsive Concurrent DNN Inference on Mobile Devices
di: Lin, Zheng, et al.
Pubblicazione: (2024)
di: Lin, Zheng, et al.
Pubblicazione: (2024)
Optimizing Distributed Training Approaches for Scaling Neural Networks
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
di: Baligodugula, Vishnu Vardhan, et al.
Pubblicazione: (2025)
Breaking the Training Barrier of Billion-Parameter Universal Machine Learning Interatomic Potentials
di: Zhou, Yuanchang, et al.
Pubblicazione: (2026)
di: Zhou, Yuanchang, et al.
Pubblicazione: (2026)
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
di: Yoon, Daegun, et al.
Pubblicazione: (2023)
di: Yoon, Daegun, et al.
Pubblicazione: (2023)
Breaking Barriers for Distributed MIS by Faster Degree Reduction
di: Khoury, Seri, et al.
Pubblicazione: (2025)
di: Khoury, Seri, et al.
Pubblicazione: (2025)
PaSE: Parallelization Strategies for Efficient DNN Training
di: Elango, Venmugil
Pubblicazione: (2024)
di: Elango, Venmugil
Pubblicazione: (2024)
Preemption Aware Task Scheduling for Priority and Deadline Constrained DNN Inference Task Offloading in Homogeneous Mobile-Edge Networks
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
di: Cotter, Jamie, et al.
Pubblicazione: (2025)
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
di: Niam, Arefin, et al.
Pubblicazione: (2025)
di: Niam, Arefin, et al.
Pubblicazione: (2025)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
di: Lian, Xinyu, et al.
Pubblicazione: (2024)
di: Lian, Xinyu, et al.
Pubblicazione: (2024)
Collaborative Satellite Computing through Adaptive DNN Task Splitting and Offloading
di: Peng, Shifeng, et al.
Pubblicazione: (2024)
di: Peng, Shifeng, et al.
Pubblicazione: (2024)
Pagoda: An Energy and Time Roofline Study for DNN Workloads on Edge Accelerators
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
di: K., Prashanthi S., et al.
Pubblicazione: (2025)
Documenti analoghi
-
Nezha: A Key-Value Separated Distributed Store with Optimized Raft Integration
di: Wang, Yangyang, et al.
Pubblicazione: (2026) -
UNR: Unified Notifiable RMA Library for HPC
di: Feng, Guangnan, et al.
Pubblicazione: (2024) -
Demystifying ARM SME to Optimize General Matrix Multiplications
di: Deng, Chencheng, et al.
Pubblicazione: (2025) -
Distributed Set-membership Filtering Frameworks For Multi-agent Systems With Absolute and Relative Measurements
di: Ding, Yu, et al.
Pubblicazione: (2023) -
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
di: Zhang, WenZheng, et al.
Pubblicazione: (2024)