AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
Fuente:
arXiv
Saved in:
| Main Authors: | Xiao, Youshao, Ju, Lin, Zhou, Zhenglei, Li, Siyuan, Huan, Zhaoxin, Zhang, Dalong, Jiang, Rujie, Wang, Lin, Zhang, Xiaolu, Liang, Lei, Zhou, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024)
by: Xiao, Youshao, et al.
Published: (2024)
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025)
by: Wu, Tianyuan, et al.
Published: (2025)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)
by: Li, Haoyang, et al.
Published: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
by: Adikari, Tharindu, et al.
Published: (2024)
by: Adikari, Tharindu, et al.
Published: (2024)
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
by: Yan, Jiahe, et al.
Published: (2024)
by: Yan, Jiahe, et al.
Published: (2024)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
by: Wu, Tianyuan, et al.
Published: (2024)
by: Wu, Tianyuan, et al.
Published: (2024)
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
by: Liu, Guanliang, et al.
Published: (2026)
by: Liu, Guanliang, et al.
Published: (2026)
Understanding Stragglers in Large Model Training Using What-if Analysis
by: Lin, Jinkun, et al.
Published: (2025)
by: Lin, Jinkun, et al.
Published: (2025)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
by: Liu, Xing, et al.
Published: (2025)
by: Liu, Xing, et al.
Published: (2025)
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
by: Das, Anindya Bijoy, et al.
Published: (2024)
by: Das, Anindya Bijoy, et al.
Published: (2024)
Self-healing Nodes with Adaptive Data-Sharding
by: Thakur, Ayush, et al.
Published: (2024)
by: Thakur, Ayush, et al.
Published: (2024)
Lightweight Federated Learning with Differential Privacy and Straggler Resilience
by: Hong, Shu, et al.
Published: (2024)
by: Hong, Shu, et al.
Published: (2024)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
by: Xiong, Guojun, et al.
Published: (2023)
by: Xiong, Guojun, et al.
Published: (2023)
DistFlow: A Fully Distributed RL Framework for Scalable and Efficient LLM Post-Training
by: Wang, Zhixin, et al.
Published: (2025)
by: Wang, Zhixin, et al.
Published: (2025)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
OpenDT: Exploring Datacenter Performance and Sustainability with a Self-Calibrating Digital Twin
by: Nicolae, Radu, et al.
Published: (2026)
by: Nicolae, Radu, et al.
Published: (2026)
Distributed Learning based on 1-Bit Gradient Coding in the Presence of Stragglers
by: Li, Chengxi, et al.
Published: (2024)
by: Li, Chengxi, et al.
Published: (2024)
DeFT: Mitigating Data Dependencies for Flexible Communication Scheduling in Distributed Training
by: Meng, Lin, et al.
Published: (2025)
by: Meng, Lin, et al.
Published: (2025)
Efficient AllReduce with Stragglers
by: Devraj, Arjun, et al.
Published: (2025)
by: Devraj, Arjun, et al.
Published: (2025)
Beyond 2-Edge-Connectivity: Algorithms and Impossibility for Content-Oblivious Leader Election
by: Chang, Yi-Jun, et al.
Published: (2025)
by: Chang, Yi-Jun, et al.
Published: (2025)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
by: Ma, Xiaolong, et al.
Published: (2024)
by: Ma, Xiaolong, et al.
Published: (2024)
LiveR: Fine-Grained Elasticity via Live Reconfiguration for Model Training
by: Liu, Haoyuan, et al.
Published: (2026)
by: Liu, Haoyuan, et al.
Published: (2026)
Efficient Training of Large Language Models on Distributed Infrastructures: A Survey
by: Duan, Jiangfei, et al.
Published: (2024)
by: Duan, Jiangfei, et al.
Published: (2024)
Content-Oblivious Leader Election in 2-Edge-Connected Networks
by: Chalopin, Jérémie, et al.
Published: (2025)
by: Chalopin, Jérémie, et al.
Published: (2025)
Silent Self-Stabilising Leader Election in Programmable Matter Systems with Holes
by: Chalopin, Jérémie, et al.
Published: (2025)
by: Chalopin, Jérémie, et al.
Published: (2025)
Deterministic Self-Stabilising Leader Election for Programmable Matter with Constant Memory
by: Chalopin, Jérémie, et al.
Published: (2024)
by: Chalopin, Jérémie, et al.
Published: (2024)
Non-Uniform Content-Oblivious Leader Election on Oriented Asynchronous Rings
by: Chalopin, Jérémie, et al.
Published: (2025)
by: Chalopin, Jérémie, et al.
Published: (2025)
Accelerating Compound LLM Training Workloads with Maestro
by: Yuan, Xiulong, et al.
Published: (2026)
by: Yuan, Xiulong, et al.
Published: (2026)
A Survey of Distributed Graph Algorithms on Massive Graphs
by: Meng, Lingkai, et al.
Published: (2024)
by: Meng, Lingkai, et al.
Published: (2024)
Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous Platforms
by: Qiao, Tong, et al.
Published: (2025)
by: Qiao, Tong, et al.
Published: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
GreenDyGNN: Runtime-Adaptive Energy-Efficient Communication for Distributed GNN Training
by: Niam, Arefin, et al.
Published: (2026)
by: Niam, Arefin, et al.
Published: (2026)
General Coded Computing in a Probabilistic Straggler Regime
by: Moradi, Parsa, et al.
Published: (2025)
by: Moradi, Parsa, et al.
Published: (2025)
A Space-Time Trade-off for Fast Self-Stabilizing Leader Election in Population Protocols
by: Austin, Henry, et al.
Published: (2025)
by: Austin, Henry, et al.
Published: (2025)
FuxiShuffle: An Adaptive and Resilient Shuffle Service for Distributed Data Processing on Alibaba Cloud
by: Lin, Yuhao, et al.
Published: (2026)
by: Lin, Yuhao, et al.
Published: (2026)
CHIRON: Accelerating Node Synchronization without Security Trade-offs in Distributed Ledgers
by: Neiheiser, Ray, et al.
Published: (2024)
by: Neiheiser, Ray, et al.
Published: (2024)
Daedalus: Self-Adaptive Horizontal Autoscaling for Resource Efficiency of Distributed Stream Processing Systems
by: Pfister, Benjamin J. J., et al.
Published: (2024)
by: Pfister, Benjamin J. J., et al.
Published: (2024)
Similar Items
-
G-Meta: Distributed Meta Learning in GPU Clusters for Large-Scale Recommender Systems
by: Xiao, Youshao, et al.
Published: (2024) -
AntBatchInfer: Elastic Batch Inference in the Kubernetes Cluster
by: Li, Siyuan, et al.
Published: (2024) -
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
by: Wu, Tianyuan, et al.
Published: (2025) -
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025) -
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
by: Li, Haoyang, et al.
Published: (2024)