Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
Fuente:
arXiv
Saved in:
| Main Authors: | He, Guoliang, Jiang, Youhe, Xiao, Wencong, Jiang, Kaihua, Wang, Shuguang, Wang, Jun, Du, Zixian, Jiang, Zhuo, Zhang, Xinlei, Yuan, Binhang, Yoneki, Eiko |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
by: Liang, Yan, et al.
Published: (2026)
by: Liang, Yan, et al.
Published: (2026)
Efficient Multi-round LLM Inference over Disaggregated Serving
by: He, Wenhao, et al.
Published: (2026)
by: He, Wenhao, et al.
Published: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
by: Yan, Ran, et al.
Published: (2024)
by: Yan, Ran, et al.
Published: (2024)
AReaL-Hex: Accommodating Asynchronous RL Training over Heterogeneous GPUs
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
by: Peng, You, et al.
Published: (2026)
by: Peng, You, et al.
Published: (2026)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2023)
by: Jiang, Youhe, et al.
Published: (2023)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
by: Jiang, Youhe, et al.
Published: (2026)
by: Jiang, Youhe, et al.
Published: (2026)
MegaScale: Scaling Large Language Model Training to More Than 10,000 GPUs
by: Jiang, Ziheng, et al.
Published: (2024)
by: Jiang, Ziheng, et al.
Published: (2024)
Cascadia: An Efficient Cascade Serving System for Large Language Models
by: Jiang, Youhe, et al.
Published: (2025)
by: Jiang, Youhe, et al.
Published: (2025)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
by: Yan, Ran, et al.
Published: (2025)
by: Yan, Ran, et al.
Published: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
by: Tong, Chris, et al.
Published: (2025)
by: Tong, Chris, et al.
Published: (2025)
Revisiting the Time Cost Model of AllReduce
by: Xiong, Dian, et al.
Published: (2024)
by: Xiong, Dian, et al.
Published: (2024)
Mycroft: Tracing Dependencies in Collective Communication Towards Reliable LLM Training
by: Deng, Yangtao, et al.
Published: (2025)
by: Deng, Yangtao, et al.
Published: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
by: Zhang, Li, et al.
Published: (2025)
by: Zhang, Li, et al.
Published: (2025)
Training LLMs with Fault Tolerant HSDP on 100,000 GPUs
by: Salpekar, Omkar, et al.
Published: (2026)
by: Salpekar, Omkar, et al.
Published: (2026)
Straggler Tolerant and Resilient DL Training on Homogeneous GPUs
by: Zhang, Zeyu, et al.
Published: (2025)
by: Zhang, Zeyu, et al.
Published: (2025)
TrioSeq: A Novel Approach to Accelerate Triplet Sequence Alignment on GPUs
by: Graça, Miguel, et al.
Published: (2026)
by: Graça, Miguel, et al.
Published: (2026)
Improving Automatic Parallel Training via Balanced Memory Workload Optimization
by: Wang, Yujie, et al.
Published: (2023)
by: Wang, Yujie, et al.
Published: (2023)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
by: Ge, Hao, et al.
Published: (2025)
by: Ge, Hao, et al.
Published: (2025)
TensorHub: Scalable and Elastic Weight Transfer for LLM RL Training
by: Ye, Chenhao, et al.
Published: (2026)
by: Ye, Chenhao, et al.
Published: (2026)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Analyzing the Performance Portability of SYCL across CPUs, GPUs, and Hybrid Systems with SW Sequence Alignment
by: Costanzo, Manuel, et al.
Published: (2024)
by: Costanzo, Manuel, et al.
Published: (2024)
Nixie: Efficient, Transparent Temporal Multiplexing for Consumer GPUs
by: Xu, Yechen, et al.
Published: (2026)
by: Xu, Yechen, et al.
Published: (2026)
DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language Models
by: Zhang, Zili, et al.
Published: (2024)
by: Zhang, Zili, et al.
Published: (2024)
HeterMoE: Efficient Training of Mixture-of-Experts Models on Heterogeneous GPUs
by: Wu, Yongji, et al.
Published: (2025)
by: Wu, Yongji, et al.
Published: (2025)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
ROSE: Rollout On Serving GPUs via Cooperative Elasticity for Agentic RL
by: Gao, Wei, et al.
Published: (2026)
by: Gao, Wei, et al.
Published: (2026)
MTGenRec: An Efficient Distributed Training System for Generative Recommendation Models in Meituan
by: Wang, Yuxiang, et al.
Published: (2025)
by: Wang, Yuxiang, et al.
Published: (2025)
Schedule-Level Shared-Prefix Reuse for LLM RL Training
by: Li, Pengbo, et al.
Published: (2026)
by: Li, Pengbo, et al.
Published: (2026)
An Adaptive Distributed Stencil Abstraction for GPUs
by: Bhosale, Aditya, et al.
Published: (2025)
by: Bhosale, Aditya, et al.
Published: (2025)
Accelerating Maximal Biclique Enumeration on GPUs
by: Hsieh, Chou-Ying, et al.
Published: (2024)
by: Hsieh, Chou-Ying, et al.
Published: (2024)
Parallelizing Maximal Clique Enumeration on GPUs
by: Almasri, Mohammad, et al.
Published: (2022)
by: Almasri, Mohammad, et al.
Published: (2022)
Optimizing sDTW for AMD GPUs
by: Latta-Lin, Daniel, et al.
Published: (2024)
by: Latta-Lin, Daniel, et al.
Published: (2024)
FastCHGNet: Training one Universal Interatomic Potential to 1.5 Hours with 32 GPUs
by: Zhou, Yuanchang, et al.
Published: (2024)
by: Zhou, Yuanchang, et al.
Published: (2024)
FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion
by: Chang, Li-Wen, et al.
Published: (2024)
by: Chang, Li-Wen, et al.
Published: (2024)
Similar Items
-
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
by: Jiang, Youhe, et al.
Published: (2026) -
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
by: Jiang, Youhe, et al.
Published: (2025) -
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
by: Jiang, Youhe, et al.
Published: (2026) -
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
by: Jiang, Youhe, et al.
Published: (2025) -
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
by: Jiang, Youhe, et al.
Published: (2025)