Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Zeyu, Shen, Haiying |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
por: Yan, Jiahe, et al.
Publicado: (2024)
por: Yan, Jiahe, et al.
Publicado: (2024)
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
por: Wu, Tianyuan, et al.
Publicado: (2025)
por: Wu, Tianyuan, et al.
Publicado: (2025)
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
por: Li, Haoyang, et al.
Publicado: (2024)
por: Li, Haoyang, et al.
Publicado: (2024)
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
por: Zhang, Zeyu, et al.
Publicado: (2024)
por: Zhang, Zeyu, et al.
Publicado: (2024)
Uncoded Download in Lagrange-Coded Elastic Computing with Straggler Tolerance
por: Zhong, Xi, et al.
Publicado: (2025)
por: Zhong, Xi, et al.
Publicado: (2025)
Lightweight Federated Learning with Differential Privacy and Straggler Resilience
por: Hong, Shu, et al.
Publicado: (2024)
por: Hong, Shu, et al.
Publicado: (2024)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
por: Salpekar, Omkar, et al.
Publicado: (2026)
por: Salpekar, Omkar, et al.
Publicado: (2026)
Straggler-Resilient Decentralized Learning via Adaptive Asynchronous Updates
por: Xiong, Guojun, et al.
Publicado: (2023)
por: Xiong, Guojun, et al.
Publicado: (2023)
EconoServe: Maximizing Multi-Resource Utilization with SLO Guarantees in LLM Serving
por: Shen, Haiying, et al.
Publicado: (2024)
por: Shen, Haiying, et al.
Publicado: (2024)
Exploiting Stragglers in Distributed Computing Systems with Task Grouping
por: Adikari, Tharindu, et al.
Publicado: (2024)
por: Adikari, Tharindu, et al.
Publicado: (2024)
TurboFFT: Co-Designed High-Performance and Fault-Tolerant Fast Fourier Transform on GPUs
por: Wu, Shixun, et al.
Publicado: (2024)
por: Wu, Shixun, et al.
Publicado: (2024)
CoCoI: Distributed Coded Inference System for Straggler Mitigation
por: Liu, Xing, et al.
Publicado: (2025)
por: Liu, Xing, et al.
Publicado: (2025)
FALCON: Pinpointing and Mitigating Stragglers for Large-Scale Hybrid-Parallel Training
por: Wu, Tianyuan, et al.
Publicado: (2024)
por: Wu, Tianyuan, et al.
Publicado: (2024)
Sparsity-Preserving Encodings for Straggler-Optimal Distributed Matrix Computations at the Edge
por: Das, Anindya Bijoy, et al.
Publicado: (2024)
por: Das, Anindya Bijoy, et al.
Publicado: (2024)
Understanding Stragglers in Large Model Training Using What-if Analysis
por: Lin, Jinkun, et al.
Publicado: (2025)
por: Lin, Jinkun, et al.
Publicado: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
por: Tanaka, Masahiro, et al.
Publicado: (2025)
por: Tanaka, Masahiro, et al.
Publicado: (2025)
AntDT: A Self-Adaptive Distributed Training Framework for Leader and Straggler Nodes
por: Xiao, Youshao, et al.
Publicado: (2024)
por: Xiao, Youshao, et al.
Publicado: (2024)
Efficient AllReduce with Stragglers
por: Devraj, Arjun, et al.
Publicado: (2025)
por: Devraj, Arjun, et al.
Publicado: (2025)
Story of Two GPUs: Characterizing the Resilience of Hopper H100 and Ampere A100 GPUs
por: Cui, Shengkun, et al.
Publicado: (2025)
por: Cui, Shengkun, et al.
Publicado: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
por: He, Guoliang, et al.
Publicado: (2025)
por: He, Guoliang, et al.
Publicado: (2025)
SpecInF: Exploiting Idle GPU Resources in Distributed DL Training via Speculative Inference Filling
por: Lv, Cunchi, et al.
Publicado: (2025)
por: Lv, Cunchi, et al.
Publicado: (2025)
ESG: Pipeline-Conscious Efficient Scheduling of DNN Workflows on Serverless Platforms with Shareable GPUs
por: Hui, Xinning, et al.
Publicado: (2024)
por: Hui, Xinning, et al.
Publicado: (2024)
Towards Straggler-Resilient Split Federated Learning: An Unbalanced Update Approach
por: Liang, Dandan, et al.
Publicado: (2025)
por: Liang, Dandan, et al.
Publicado: (2025)
An Adaptive Distributed Stencil Abstraction for GPUs
por: Bhosale, Aditya, et al.
Publicado: (2025)
por: Bhosale, Aditya, et al.
Publicado: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
por: Hsieh, Chou-Ying, et al.
Publicado: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
por: Almasri, Mohammad, et al.
Publicado: (2022)
por: Almasri, Mohammad, et al.
Publicado: (2022)
Optimizing sDTW for AMD GPUs
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
por: Latta-Lin, Daniel, et al.
Publicado: (2024)
Uncoded Storage Coded Transmission Elastic Computing with Straggler Tolerance in Heterogeneous Systems
por: Zhong, Xi, et al.
Publicado: (2024)
por: Zhong, Xi, et al.
Publicado: (2024)
Serving Compound Inference Systems on Datacenter GPUs
por: Devata, Sriram, et al.
Publicado: (2026)
por: Devata, Sriram, et al.
Publicado: (2026)
Fast Kronecker Matrix-Matrix Multiplication on GPUs
por: Jangda, Abhinav, et al.
Publicado: (2024)
por: Jangda, Abhinav, et al.
Publicado: (2024)
Optimal Workload Placement on Multi-Instance GPUs
por: Turkkan, Bekir, et al.
Publicado: (2024)
por: Turkkan, Bekir, et al.
Publicado: (2024)
FlashMP: Fast Discrete Transform-Based Solver for Preconditioning Maxwell's Equations on GPUs
por: Zhang, Haoyuan, et al.
Publicado: (2025)
por: Zhang, Haoyuan, et al.
Publicado: (2025)
Chameleon: Adaptive Fault Tolerance for Distributed Training via Real-time Policy Selection
por: Zhou, Yuhang, et al.
Publicado: (2025)
por: Zhou, Yuhang, et al.
Publicado: (2025)
HACK: Homomorphic Acceleration via Compression of the Key-Value Cache for Disaggregated LLM Inference
por: Zhang, Zeyu, et al.
Publicado: (2025)
por: Zhang, Zeyu, et al.
Publicado: (2025)
RDMA-Based Algorithms for Sparse Matrix Multiplication on GPUs
por: Brock, Benjamin, et al.
Publicado: (2023)
por: Brock, Benjamin, et al.
Publicado: (2023)
Accurate Computation of the Logarithm of Modified Bessel Functions on GPUs
por: Plesner, Andreas, et al.
Publicado: (2024)
por: Plesner, Andreas, et al.
Publicado: (2024)
ReCoVer: Resilient LLM Pre-Training System via Fault-Tolerant Collective and Versatile Workload
por: Liu, Ziyue, et al.
Publicado: (2026)
por: Liu, Ziyue, et al.
Publicado: (2026)
General Coded Computing in a Probabilistic Straggler Regime
por: Moradi, Parsa, et al.
Publicado: (2025)
por: Moradi, Parsa, et al.
Publicado: (2025)
SPARe: Stacked Parallelism with Adaptive Reordering for Fault-Tolerant LLM Pretraining Systems with 100k+ GPUs
por: Lee, Jin, et al.
Publicado: (2026)
por: Lee, Jin, et al.
Publicado: (2026)
Ejemplares similares
-
dSTAR: Straggler Tolerant and Byzantine Resilient Distributed SGD
por: Yan, Jiahe, et al.
Publicado: (2024) -
Adaptra: Straggler-Resilient Hybrid-Parallel Training with Pipeline Adaptation
por: Wu, Tianyuan, et al.
Publicado: (2025) -
Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model Parallelization
por: Li, Haoyang, et al.
Publicado: (2024) -
FDC: Fast KV Dimensionality Compression for Efficient LLM Inference
por: Zhang, Zeyu, et al.
Publicado: (2024) -
PecSched: Preemptive and Efficient Cluster Scheduling for LLM Inference
por: Zhang, Zeyu, et al.
Publicado: (2024)