Revisiting the Time Cost Model of AllReduce
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xiong, Dian, Chen, Li, Jiang, Youhe, Li, Dan, Wang, Shuai, Wang, Songtao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient AllReduce with Stragglers
von: Devraj, Arjun, et al.
Veröffentlicht: (2025)
von: Devraj, Arjun, et al.
Veröffentlicht: (2025)
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
von: Chen, Peiqing, et al.
Veröffentlicht: (2026)
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
von: Wei, Yufan, et al.
Veröffentlicht: (2025)
Short-circuiting Rings for Low-Latency AllReduce
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
von: Juerss, Anton, et al.
Veröffentlicht: (2026)
BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
Rina: Enhancing Ring-AllReduce with In-network Aggregation in Distributed Model Training
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
von: Chen, Zixuan, et al.
Veröffentlicht: (2024)
OptiReduce: Resilient and Tail-Optimal AllReduce for Distributed Deep Learning in the Cloud
von: Warraich, Ertza, et al.
Veröffentlicht: (2023)
von: Warraich, Ertza, et al.
Veröffentlicht: (2023)
ThunderServe: High-performance and Cost-efficient LLM Serving in Cloud Environments
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
HexGen-2: Disaggregated Generative Inference of LLMs in Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
OServe: Accelerating LLM Serving via Spatial-Temporal Workload Orchestration
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
HexAGenT: Efficient Agentic LLM Serving via Workflow- and Heterogeneity-Aware Scheduling
von: Peng, You, et al.
Veröffentlicht: (2026)
von: Peng, You, et al.
Veröffentlicht: (2026)
HexGen: Generative Inference of Large Language Model over Heterogeneous Environment
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
von: Jiang, Youhe, et al.
Veröffentlicht: (2023)
Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
von: Jiang, Youhe, et al.
Veröffentlicht: (2026)
HexiSeq: Accommodating Long Context Training of LLMs over Heterogeneous Hardware
von: Liang, Yan, et al.
Veröffentlicht: (2026)
von: Liang, Yan, et al.
Veröffentlicht: (2026)
HexiScale: Facilitating Large Language Model Training over Heterogeneous Hardware
von: Yan, Ran, et al.
Veröffentlicht: (2024)
von: Yan, Ran, et al.
Veröffentlicht: (2024)
Communication-Efficient Serving for Video Diffusion Models with Latent Parallelism
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Wu, Zhiyuan, et al.
Veröffentlicht: (2025)
Cascadia: An Efficient Cascade Serving System for Large Language Models
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
von: Jiang, Youhe, et al.
Veröffentlicht: (2025)
Efficient Pre-Training of LLMs via Topology-Aware Communication Alignment on More Than 9600 GPUs
von: He, Guoliang, et al.
Veröffentlicht: (2025)
von: He, Guoliang, et al.
Veröffentlicht: (2025)
Efficient Multi-round LLM Inference over Disaggregated Serving
von: He, Wenhao, et al.
Veröffentlicht: (2026)
von: He, Wenhao, et al.
Veröffentlicht: (2026)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
von: Yao, Chenxuan, et al.
Veröffentlicht: (2025)
von: Yao, Chenxuan, et al.
Veröffentlicht: (2025)
Parallax: Efficient LLM Inference Service over Decentralized Environment
von: Tong, Chris, et al.
Veröffentlicht: (2025)
von: Tong, Chris, et al.
Veröffentlicht: (2025)
LMDeploy Accelerates Mixed-Precision LLM Inference with TurboMind
von: Zhang, Li, et al.
Veröffentlicht: (2025)
von: Zhang, Li, et al.
Veröffentlicht: (2025)
Polynomial Time Local Decision Revisited
von: Feuilloley, Laurent, et al.
Veröffentlicht: (2026)
von: Feuilloley, Laurent, et al.
Veröffentlicht: (2026)
VSS Challenge Problem: Verifying the Correctness of AllReduce Algorithms in the MPICH Implementation of MPI
von: Hovland, Paul D.
Veröffentlicht: (2025)
von: Hovland, Paul D.
Veröffentlicht: (2025)
Exploiting the Uncertainty of the Longest Paths: Response Time Analysis for Probabilistic DAG Tasks
von: Gao, Yiyang, et al.
Veröffentlicht: (2025)
von: Gao, Yiyang, et al.
Veröffentlicht: (2025)
AMSP: Reducing Communication Overhead of ZeRO for Efficient LLM Training
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
von: Chen, Qiaoling, et al.
Veröffentlicht: (2023)
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
von: Yan, Ran, et al.
Veröffentlicht: (2025)
von: Yan, Ran, et al.
Veröffentlicht: (2025)
Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and Splitting
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixin, et al.
Veröffentlicht: (2024)
Optimal Broadcast Schedules in Logarithmic Time with Applications to Broadcast, All-Broadcast, Reduction and All-Reduction
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
von: Träff, Jesper Larsson
Veröffentlicht: (2024)
FedCostAware: Enabling Cost-Aware Federated Learning on the Cloud
von: Sinha, Aditya, et al.
Veröffentlicht: (2025)
von: Sinha, Aditya, et al.
Veröffentlicht: (2025)
SLO-Aware Scheduling for Large Language Model Inferences
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
von: Huang, Jinqi, et al.
Veröffentlicht: (2025)
PICO: Accelerating All k-Core Paradigms on GPU
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
von: Zhao, Chen, et al.
Veröffentlicht: (2024)
An All-Reduce Compatible Top-K Compressor for Communication-Efficient Distributed Learning
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
von: Chen, Chuyan, et al.
Veröffentlicht: (2025)
Hiding Communication Cost in Distributed LLM Training via Micro-batch Co-execution
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
von: Wang, Haiquan, et al.
Veröffentlicht: (2024)
LoHan: Low-Cost High-Performance Framework to Fine-Tune 100B Model on a Consumer GPU
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
von: Liao, Changyue, et al.
Veröffentlicht: (2024)
FaaSMT: Lightweight Serverless Framework for Intrusion Detection Using Merkle Tree and Task Inlining
von: Li, Chuang, et al.
Veröffentlicht: (2025)
von: Li, Chuang, et al.
Veröffentlicht: (2025)
Decentralized Semantic Federated Learning for Real-Time Public Safety Tasks: Challenges, Methods, and Directions
von: Li, Baosheng, et al.
Veröffentlicht: (2025)
von: Li, Baosheng, et al.
Veröffentlicht: (2025)
Large-Scale Metric Computation in Online Controlled Experiment Platform
von: Xiong, Tao, et al.
Veröffentlicht: (2024)
von: Xiong, Tao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient AllReduce with Stragglers
von: Devraj, Arjun, et al.
Veröffentlicht: (2025) -
Don't Let a Few Network Failures Slow the Entire AllReduce
von: Chen, Peiqing, et al.
Veröffentlicht: (2026) -
AllReduce Scheduling with Hierarchical Deep Reinforcement Learning
von: Wei, Yufan, et al.
Veröffentlicht: (2025) -
Short-circuiting Rings for Low-Latency AllReduce
von: Hammer, Sarah-Michelle, et al.
Veröffentlicht: (2025) -
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
von: Juerss, Anton, et al.
Veröffentlicht: (2026)