Optimizing Distributed Training Approaches for Scaling Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Baligodugula, Vishnu Vardhan, Amsaad, Fathi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025)
by: Niam, Arefin, et al.
Published: (2025)
Heta: Distributed Training of Heterogeneous Graph Neural Networks
by: Zhong, Yuchen, et al.
Published: (2024)
by: Zhong, Yuchen, et al.
Published: (2024)
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024)
by: Sun, Mo, et al.
Published: (2024)
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025)
by: Tanaka, Masahiro, et al.
Published: (2025)
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)
by: Waleffe, Roger, et al.
Published: (2025)
BANG: Billion-Scale Approximate Nearest Neighbor Search using a Single GPU
by: V., Karthik, et al.
Published: (2024)
by: V., Karthik, et al.
Published: (2024)
MalleTrain: Deep Neural Network Training on Unfillable Supercomputer Nodes
by: Ma, Xiaolong, et al.
Published: (2024)
by: Ma, Xiaolong, et al.
Published: (2024)
Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU Clusters
by: Zhang, WenZheng, et al.
Published: (2024)
by: Zhang, WenZheng, et al.
Published: (2024)
CaPGNN: Optimizing Parallel Graph Neural Network Training with Joint Caching and Resource-Aware Graph Partitioning
by: Song, Xianfeng, et al.
Published: (2025)
by: Song, Xianfeng, et al.
Published: (2025)
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training
by: Golden, Alicia, et al.
Published: (2025)
by: Golden, Alicia, et al.
Published: (2025)
Optimizing Frequent Checkpointing via Low-Cost Differential for Distributed Training Systems
by: Yao, Chenxuan, et al.
Published: (2025)
by: Yao, Chenxuan, et al.
Published: (2025)
Nezha: Breaking Multi-Rail Network Barriers for Distributed DNN Training
by: Yu, Enda, et al.
Published: (2024)
by: Yu, Enda, et al.
Published: (2024)
Effectiveness of Distributed Gradient Descent with Local Steps for Overparameterized Models
by: Zhu, Heng, et al.
Published: (2024)
by: Zhu, Heng, et al.
Published: (2024)
Optimizing High-Throughput Distributed Data Pipelines for Reproducible Deep Learning at Scale
by: Mittal, Kashish, et al.
Published: (2026)
by: Mittal, Kashish, et al.
Published: (2026)
Distributed Generative Inference of LLM at Internet Scales with Multi-Dimensional Communication Optimization
by: Chen, Jiu, et al.
Published: (2026)
by: Chen, Jiu, et al.
Published: (2026)
Multi-Resolution Model Fusion for Accelerating the Convolutional Neural Network Training
by: Wang, Kewei, et al.
Published: (2025)
by: Wang, Kewei, et al.
Published: (2025)
Fully Distributed Online Training of Graph Neural Networks in Networked Systems
by: Olshevskyi, Rostyslav, et al.
Published: (2024)
by: Olshevskyi, Rostyslav, et al.
Published: (2024)
Embedded Distributed Inference of Deep Neural Networks: A Systematic Review
by: Peccia, Federico Nicolás, et al.
Published: (2024)
by: Peccia, Federico Nicolás, et al.
Published: (2024)
Diagonal Scaling: A Multi-Dimensional Resource Model and Optimization Framework for Distributed Databases
by: Abdullah, Shahir, et al.
Published: (2025)
by: Abdullah, Shahir, et al.
Published: (2025)
ReInc: Scaling Training of Dynamic Graph Neural Networks
by: Guan, Mingyu, et al.
Published: (2025)
by: Guan, Mingyu, et al.
Published: (2025)
Distributed Convolutional Neural Network Training on Mobile and Edge Clusters
by: Rama, Pranav, et al.
Published: (2024)
by: Rama, Pranav, et al.
Published: (2024)
ACE-Sync: An Adaptive Cloud-Edge Synchronization Framework for Communication-Efficient Large-Scale Distributed Model Training
by: Yang, Yi, et al.
Published: (2025)
by: Yang, Yi, et al.
Published: (2025)
ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural Networks
by: Naman, Pranjal, et al.
Published: (2026)
by: Naman, Pranjal, et al.
Published: (2026)
Spatiotemporal Traffic Prediction in Distributed Backend Systems via Graph Neural Networks
by: Qiu, Zhimin, et al.
Published: (2025)
by: Qiu, Zhimin, et al.
Published: (2025)
A Distributed Approach for Persistent Homology Computation on a Large Scale
by: Ceccaroni, Riccardo, et al.
Published: (2024)
by: Ceccaroni, Riccardo, et al.
Published: (2024)
Efficient Distributed MLLM Training with Cornstarch
by: Jang, Insu, et al.
Published: (2025)
by: Jang, Insu, et al.
Published: (2025)
An Experimental Comparison of Partitioning Strategies for Distributed Graph Neural Network Training
by: Merkel, Nikolai, et al.
Published: (2023)
by: Merkel, Nikolai, et al.
Published: (2023)
NEST: Network- and Memory-Aware Device Placement For Distributed Deep Learning
by: Wang, Irene, et al.
Published: (2026)
by: Wang, Irene, et al.
Published: (2026)
Distributed Constrained Combinatorial Optimization leveraging Hypergraph Neural Networks
by: Heydaribeni, Nasimeh, et al.
Published: (2023)
by: Heydaribeni, Nasimeh, et al.
Published: (2023)
Echo: Simulating Distributed Training At Scale
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
Stable-MoE: Lyapunov-based Token Routing for Distributed Mixture-of-Experts Training over Edge Networks
by: Shi, Long, et al.
Published: (2025)
by: Shi, Long, et al.
Published: (2025)
OptimES: Optimizing Federated Learning Using Remote Embeddings for Graph Neural Networks
by: Naman, Pranjal, et al.
Published: (2025)
by: Naman, Pranjal, et al.
Published: (2025)
SDT-GNN: Streaming-based Distributed Training Framework for Graph Neural Networks
by: Huang, Xin, et al.
Published: (2024)
by: Huang, Xin, et al.
Published: (2024)
Galvatron: Automatic Distributed Training for Large Transformer Models
by: Gumaan, Esmail
Published: (2025)
by: Gumaan, Esmail
Published: (2025)
Addressing Variable Heterogeneity in Distributed Multimodal Training with Entrain
by: Jang, Insu, et al.
Published: (2026)
by: Jang, Insu, et al.
Published: (2026)
Accelerating Distributed MoE Training and Inference with Lina
by: Li, Jiamin, et al.
Published: (2022)
by: Li, Jiamin, et al.
Published: (2022)
Collaborative UAVs Multi-task Video Processing Optimization Based on Enhanced Distributed Actor-Critic Networks
by: Rong, Ziqi, et al.
Published: (2024)
by: Rong, Ziqi, et al.
Published: (2024)
GSplit: Scaling Graph Neural Network Training on Large Graphs via Split-Parallelism
by: Polisetty, Sandeep, et al.
Published: (2023)
by: Polisetty, Sandeep, et al.
Published: (2023)
Federated Neural Radiance Field for Distributed Intelligence
by: Zhang, Yintian, et al.
Published: (2024)
by: Zhang, Yintian, et al.
Published: (2024)
MegatronApp: Efficient and Comprehensive Management on Distributed LLM Training
by: Zhao, Bohan, et al.
Published: (2025)
by: Zhao, Bohan, et al.
Published: (2025)
Similar Items
-
RapidGNN: Communication Efficient Large-Scale Distributed Training of Graph Neural Networks
by: Niam, Arefin, et al.
Published: (2025) -
Heta: Distributed Training of Heterogeneous Graph Neural Networks
by: Zhong, Yuchen, et al.
Published: (2024) -
LuWu: An End-to-End In-Network Out-of-Core Optimizer for 100B-Scale Model-in-Network Data-Parallel Training on Distributed GPUs
by: Sun, Mo, et al.
Published: (2024) -
DeepCompile: A Compiler-Driven Approach to Optimizing Distributed Deep Learning Training
by: Tanaka, Masahiro, et al.
Published: (2025) -
Armada: Memory-Efficient Distributed Training of Large-Scale Graph Neural Networks
by: Waleffe, Roger, et al.
Published: (2025)