The Big Send-off: Scalable and Performant Collectives for Deep Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Singh, Siddharth, Pradeep, Keshav, Singh, Mahua, Wei, Cunyang, Bhatele, Abhinav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025)
by: Ranjan, Aditya K., et al.
Published: (2025)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
by: Wei, Cunyang, et al.
Published: (2026)
by: Wei, Cunyang, et al.
Published: (2026)
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023)
by: Singh, Siddharth, et al.
Published: (2023)
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)
by: Chaturvedi, Aman, et al.
Published: (2024)
Understanding and Improving Communication Performance in Multi-node LLM Inference
by: Singhania, Prajwal, et al.
Published: (2025)
by: Singhania, Prajwal, et al.
Published: (2025)
Analytics of Longitudinal System Monitoring Data for Performance Prediction
by: Costello, Ian J., et al.
Published: (2020)
by: Costello, Ian J., et al.
Published: (2020)
Efficient Fine-Grained GPU Performance Modeling for Distributed Deep Learning of LLM
by: Zhang, Biyao, et al.
Published: (2025)
by: Zhang, Biyao, et al.
Published: (2025)
Tenplex: Dynamic Parallelism for Deep Learning using Parallelizable Tensor Collections
by: Wagenländer, Marcel, et al.
Published: (2023)
by: Wagenländer, Marcel, et al.
Published: (2023)
HPC-Coder: Modeling Parallel Programs using Large Language Models
by: Nichols, Daniel, et al.
Published: (2023)
by: Nichols, Daniel, et al.
Published: (2023)
Can Large Language Models Write Parallel Code?
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
SplitZip: Ultra Fast Lossless KV Compression for Disaggregated LLM Serving
by: Guo, Yipin, et al.
Published: (2026)
by: Guo, Yipin, et al.
Published: (2026)
Hubs and Spokes Learning: Efficient and Scalable Collaborative Machine Learning
by: Sharma, Atul, et al.
Published: (2025)
by: Sharma, Atul, et al.
Published: (2025)
GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline Parallelism
by: Jeon, Byungsoo, et al.
Published: (2024)
by: Jeon, Byungsoo, et al.
Published: (2024)
Scalable Explainability-as-a-Service (XaaS) for Edge AI Systems
by: Singh, Samaresh Kumar, et al.
Published: (2026)
by: Singh, Samaresh Kumar, et al.
Published: (2026)
Efficient Resource Scheduling for Distributed Infrastructures Using Negotiation Capabilities
by: Chu, Junjie, et al.
Published: (2024)
by: Chu, Junjie, et al.
Published: (2024)
FedCGD: Collective Gradient Divergence Optimized Scheduling for Wireless Federated Learning
by: Chen, Tan, et al.
Published: (2025)
by: Chen, Tan, et al.
Published: (2025)
Performance-Aligned LLMs for Generating Fast Code
by: Nichols, Daniel, et al.
Published: (2024)
by: Nichols, Daniel, et al.
Published: (2024)
Empirical Analysis of Asynchronous Federated Learning on Heterogeneous Devices: Efficiency, Fairness, and Privacy Trade-offs
by: Mohammadi, Samaneh, et al.
Published: (2025)
by: Mohammadi, Samaneh, et al.
Published: (2025)
Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs
by: Pan, Rui, et al.
Published: (2025)
by: Pan, Rui, et al.
Published: (2025)
Sentinel: Dynamic Knowledge Distillation for Personalized Federated Intrusion Detection in Heterogeneous IoT Networks
by: Singh, Gurpreet, et al.
Published: (2025)
by: Singh, Gurpreet, et al.
Published: (2025)
Deep Reinforcement Learning for System-on-Chip: Myths and Realities
by: Sung, Tegg Taekyong, et al.
Published: (2022)
by: Sung, Tegg Taekyong, et al.
Published: (2022)
Interpretable Modeling of Deep Reinforcement Learning Driven Scheduling
by: Li, Boyang, et al.
Published: (2024)
by: Li, Boyang, et al.
Published: (2024)
EdgeRL: Reinforcement Learning-driven Deep Learning Model Inference Optimization at Edge
by: Mounesan, Motahare, et al.
Published: (2024)
by: Mounesan, Motahare, et al.
Published: (2024)
Efficient and Scalable Agentic AI with Heterogeneous Systems
by: Asgar, Zain, et al.
Published: (2025)
by: Asgar, Zain, et al.
Published: (2025)
Context Parallelism for Scalable Million-Token Inference
by: Yang, Amy, et al.
Published: (2024)
by: Yang, Amy, et al.
Published: (2024)
Scale-up Unlearnable Examples Learning with High-Performance Computing
by: Zhu, Yanfan, et al.
Published: (2025)
by: Zhu, Yanfan, et al.
Published: (2025)
Scalable Artificial Intelligence for Science: Perspectives, Methods and Exemplars
by: Brewer, Wesley, et al.
Published: (2024)
by: Brewer, Wesley, et al.
Published: (2024)
Deep Reinforcement Learning for Optimizing Energy Consumption in Smart Grid Systems
by: Alsheikhi, Abeer, et al.
Published: (2026)
by: Alsheikhi, Abeer, et al.
Published: (2026)
Scaling Deep Learning Research with Kubernetes on the NRP Nautilus HyperCluster
by: Hurt, J. Alex, et al.
Published: (2024)
by: Hurt, J. Alex, et al.
Published: (2024)
Context-Aware Inference via Performance Forecasting in Decentralized Learning Networks
by: Pfeffer, Joel, et al.
Published: (2025)
by: Pfeffer, Joel, et al.
Published: (2025)
PipeOffload: Improving Scalability of Pipeline Parallelism with Memory Optimization
by: Wan, Xinyi, et al.
Published: (2025)
by: Wan, Xinyi, et al.
Published: (2025)
Laminar: A Scalable Asynchronous RL Post-Training Framework
by: Sheng, Guangming, et al.
Published: (2025)
by: Sheng, Guangming, et al.
Published: (2025)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
DistShap: Scalable GNN Explanations with Distributed Shapley Values
by: Akkas, Selahattin, et al.
Published: (2025)
by: Akkas, Selahattin, et al.
Published: (2025)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey
by: Liu, Zhihong, et al.
Published: (2024)
by: Liu, Zhihong, et al.
Published: (2024)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
by: Oakley, Joe, et al.
Published: (2024)
by: Oakley, Joe, et al.
Published: (2024)
Research on Edge Computing and Cloud Collaborative Resource Scheduling Optimization Based on Deep Reinforcement Learning
by: Wang, Yuqing, et al.
Published: (2025)
by: Wang, Yuqing, et al.
Published: (2025)
Scalable Pretraining of Large Mixture of Experts Language Models on Aurora Super Computer
by: Vooturi, Dharma Teja, et al.
Published: (2026)
by: Vooturi, Dharma Teja, et al.
Published: (2026)
Similar Items
-
Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN Training
by: Ranjan, Aditya K., et al.
Published: (2025) -
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
by: Wei, Cunyang, et al.
Published: (2026) -
A 4D Hybrid Algorithm to Scale Parallel Training to Thousands of GPUs
by: Singh, Siddharth, et al.
Published: (2023) -
Democratizing AI: Open-source Scalable LLM Training on GPU-based Supercomputers
by: Singh, Siddharth, et al.
Published: (2025) -
HPC-Coder-V2: Studying Code LLMs Across Low-Resource Parallel Languages
by: Chaturvedi, Aman, et al.
Published: (2024)