Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Guanhua, Zhang, Chengming, Shen, Zheyu, Li, Ang, Ruwase, Olatunji |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
von: Lian, Xinyu, et al.
Veröffentlicht: (2025)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
von: Shen, Zheyu, et al.
Veröffentlicht: (2025)
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)
Cost-Efficient LLM Training with Lifetime-Aware Tensor Offloading via GPUDirect Storage
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
von: Yuan, Ziqi, et al.
Veröffentlicht: (2025)
Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
von: Lian, Xinyu, et al.
Veröffentlicht: (2024)
MoEless: Efficient MoE LLM Serving via Serverless Computing
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
von: Yu, Hanfei, et al.
Veröffentlicht: (2026)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
TAPAS: Fast and Automatic Derivation of Tensor Parallel Strategies for Large Neural Networks
von: Shi, Ziji, et al.
Veröffentlicht: (2023)
von: Shi, Ziji, et al.
Veröffentlicht: (2023)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
von: Lu, Yunchi, et al.
Veröffentlicht: (2025)
von: Lu, Yunchi, et al.
Veröffentlicht: (2025)
A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
von: Li, Xiaocan, et al.
Veröffentlicht: (2025)
von: Li, Xiaocan, et al.
Veröffentlicht: (2025)
SmartMem: Layout Transformation Elimination and Adaptation for Efficient DNN Execution on Mobile
von: Niu, Wei, et al.
Veröffentlicht: (2024)
von: Niu, Wei, et al.
Veröffentlicht: (2024)
RL in the Wild: Characterizing RLVR Training in LLM Deployment
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
von: Zhou, Jiecheng, et al.
Veröffentlicht: (2025)
TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM Training
von: Liu, Man, et al.
Veröffentlicht: (2026)
von: Liu, Man, et al.
Veröffentlicht: (2026)
Robust LLM Training Infrastructure at ByteDance
von: Wan, Borui, et al.
Veröffentlicht: (2025)
von: Wan, Borui, et al.
Veröffentlicht: (2025)
Comet: Fine-grained Computation-communication Overlapping for Mixture-of-Experts
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
von: Zhang, Shulai, et al.
Veröffentlicht: (2025)
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
von: Dong, Jianbo, et al.
Veröffentlicht: (2024)
von: Dong, Jianbo, et al.
Veröffentlicht: (2024)
ProTrain: Efficient LLM Training via Memory-Aware Techniques
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
von: Yang, Hanmei, et al.
Veröffentlicht: (2024)
Role-Based Fault Tolerance System for LLM RL Post-Training
von: Chen, Zhenqian, et al.
Veröffentlicht: (2025)
von: Chen, Zhenqian, et al.
Veröffentlicht: (2025)
Distributed Low-Communication Training with Decoupled Momentum Optimization
von: Nedelkoski, Sasho, et al.
Veröffentlicht: (2025)
von: Nedelkoski, Sasho, et al.
Veröffentlicht: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
von: Tang, Zhenheng, et al.
Veröffentlicht: (2024)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
von: Li, Rui, et al.
Veröffentlicht: (2025)
von: Li, Rui, et al.
Veröffentlicht: (2025)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
von: Tanaka, Masahiro, et al.
Veröffentlicht: (2025)
Trust-free Personalized Decentralized Learning
von: Li, Yawen, et al.
Veröffentlicht: (2024)
von: Li, Yawen, et al.
Veröffentlicht: (2024)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
von: Singh, Siddharth, et al.
Veröffentlicht: (2025)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
von: Coquelin, Daniel, et al.
Veröffentlicht: (2024)
von: Coquelin, Daniel, et al.
Veröffentlicht: (2024)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
von: Yi, Kai, et al.
Veröffentlicht: (2024)
von: Yi, Kai, et al.
Veröffentlicht: (2024)
Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
von: Jiang, Chenyu, et al.
Veröffentlicht: (2024)
Autellix: An Efficient Serving Engine for LLM Agents as General Programs
von: Luo, Michael, et al.
Veröffentlicht: (2025)
von: Luo, Michael, et al.
Veröffentlicht: (2025)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
von: Ye, Chenhao, et al.
Veröffentlicht: (2026)
Frontier: Simulating the Next Generation of LLM Inference Systems
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
von: Feng, Yicheng, et al.
Veröffentlicht: (2025)
MoEBlaze: Breaking the Memory Wall for Efficient MoE Training on Modern GPUs
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Jiyuan, et al.
Veröffentlicht: (2026)
SDP4Bit: Toward 4-bit Communication Quantization in Sharded Data Parallelism for LLM Training
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
von: Jia, Jinda, et al.
Veröffentlicht: (2024)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
von: Chen, Yanxi, et al.
Veröffentlicht: (2023)
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2023)
von: Wagenländer, Marcel, et al.
Veröffentlicht: (2023)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
von: Gao, Wei, et al.
Veröffentlicht: (2025)
von: Gao, Wei, et al.
Veröffentlicht: (2025)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
von: Kim, Han-Byul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
FastPersist: Accelerating Model Checkpointing in Deep Learning
von: Wang, Guanhua, et al.
Veröffentlicht: (2024) -
SuperOffload: Unleashing the Power of Large-Scale LLM Training on Superchips
von: Lian, Xinyu, et al.
Veröffentlicht: (2025) -
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024) -
EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices
von: Shen, Zheyu, et al.
Veröffentlicht: (2025) -
AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
von: Gupta, Ahan, et al.
Veröffentlicht: (2026)