Protocol Models: Scaling Decentralized Training with Communication-Efficient Model Parallelism
Fuente:
arXiv
Saved in:
| Main Authors: | Ramasinghe, Sameera, Ajanthan, Thalaiyasingam, Avraham, Gil, Zuo, Yan, Long, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
SENTINEL: Stagewise Integrity Verification for Pipeline Parallel Decentralized Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026)
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026)
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
by: Duan, Yiqun, et al.
Published: (2025)
by: Duan, Yiqun, et al.
Published: (2025)
Guiding Neural Collapse: Optimising Towards the Nearest Simplex Equiangular Tight Frame
by: Markou, Evan, et al.
Published: (2024)
by: Markou, Evan, et al.
Published: (2024)
Self-Supervision Improves Diffusion Models for Tabular Data Imputation
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Sharper Convergence Rates for Nonconvex Optimisation via Reduction Mappings
by: Markou, Evan, et al.
Published: (2025)
by: Markou, Evan, et al.
Published: (2025)
From Activation to Initialization: Scaling Insights for Optimizing Neural Fields
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Learning Visual Hierarchies in Hyperbolic Space for Image Retrieval
by: Wang, Ziwei, et al.
Published: (2024)
by: Wang, Ziwei, et al.
Published: (2024)
A Sampling Theory Perspective on Activations for Implicit Neural Representations
by: Saratchandran, Hemanth, et al.
Published: (2024)
by: Saratchandran, Hemanth, et al.
Published: (2024)
Protocol Learning, Decentralized Frontier Risk and the No-Off Problem
by: Long, Alexander
Published: (2024)
by: Long, Alexander
Published: (2024)
ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
by: Yang, Xianghui, et al.
Published: (2024)
by: Yang, Xianghui, et al.
Published: (2024)
Efficient Parallelization Layouts for Large-Scale Distributed Model Training
by: Hagemann, Johannes, et al.
Published: (2023)
by: Hagemann, Johannes, et al.
Published: (2023)
MoE Parallel Folding: Heterogeneous Parallelism Mappings for Efficient Large-Scale MoE Model Training with Megatron Core
by: Liu, Dennis, et al.
Published: (2025)
by: Liu, Dennis, et al.
Published: (2025)
Scaling Up Data Parallelism in Decentralized Deep Learning
by: Xie, Bing, et al.
Published: (2025)
by: Xie, Bing, et al.
Published: (2025)
MSPT: Efficient Large-Scale Physical Modeling via Parallelized Multi-Scale Attention
by: Curvo, Pedro M. P., et al.
Published: (2025)
by: Curvo, Pedro M. P., et al.
Published: (2025)
An Efficient Training Algorithm for Models with Block-wise Sparsity
by: Zhu, Ding, et al.
Published: (2025)
by: Zhu, Ding, et al.
Published: (2025)
Towards Decentralized and Sustainable Foundation Model Training with the Edge
by: Xue, Leyang, et al.
Published: (2025)
by: Xue, Leyang, et al.
Published: (2025)
Communication-Efficient Training Workload Balancing for Decentralized Multi-Agent Learning
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
by: Mohammadabadi, Seyed Mahmoud Sajjadi, et al.
Published: (2024)
DHP: Efficient Scaling of MLLM Training with Dynamic Hybrid Parallelism
by: Niu, Yifan, et al.
Published: (2026)
by: Niu, Yifan, et al.
Published: (2026)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
MegaScale-MoE: Large-Scale Communication-Efficient Training of Mixture-of-Experts Models in Production
by: Jin, Chao, et al.
Published: (2025)
by: Jin, Chao, et al.
Published: (2025)
Enhancing Parallelism in Decentralized Stochastic Convex Optimization
by: Eisen, Ofri, et al.
Published: (2025)
by: Eisen, Ofri, et al.
Published: (2025)
Parallel Scaling Law for Language Models
by: Chen, Mouxiang, et al.
Published: (2025)
by: Chen, Mouxiang, et al.
Published: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
Predictive Scaling Laws for Efficient GRPO Training of Large Reasoning Models
by: Nimmaturi, Datta, et al.
Published: (2025)
by: Nimmaturi, Datta, et al.
Published: (2025)
AMDP: Asynchronous Multi-Directional Pipeline Parallelism for Large-Scale Models Training
by: Chen, Ling, et al.
Published: (2026)
by: Chen, Ling, et al.
Published: (2026)
CELLM: An Efficient Communication in Large Language Models Training for Federated Learning
by: Vavekanand, Raja, et al.
Published: (2024)
by: Vavekanand, Raja, et al.
Published: (2024)
Accelerate Model Parallel Training by Using Efficient Graph Traversal Order in Device Placement
by: Wang, Tianze, et al.
Published: (2022)
by: Wang, Tianze, et al.
Published: (2022)
On Optimizing the Communication of Model Parallelism
by: Zhuang, Yonghao, et al.
Published: (2022)
by: Zhuang, Yonghao, et al.
Published: (2022)
Arena: Efficiently Training Large Models via Dynamic Scheduling and Adaptive Parallelism Co-Design
by: Xue, Chunyu, et al.
Published: (2024)
by: Xue, Chunyu, et al.
Published: (2024)
Jigsaw: Training Multi-Billion-Parameter AI Weather Models with Optimized Model Parallelism
by: Kieckhefen, Deifilia, et al.
Published: (2025)
by: Kieckhefen, Deifilia, et al.
Published: (2025)
Two-dimensional Sparse Parallelism for Large Scale Deep Learning Recommendation Model Training
by: Zhang, Xin, et al.
Published: (2025)
by: Zhang, Xin, et al.
Published: (2025)
Communication-Efficient Language Model Training Scales Reliably and Robustly: Scaling Laws for DiLoCo
by: Charles, Zachary, et al.
Published: (2025)
by: Charles, Zachary, et al.
Published: (2025)
Beyond Centralization: Provable Communication Efficient Decentralized Multi-Task Learning
by: Kang, Donghwa, et al.
Published: (2025)
by: Kang, Donghwa, et al.
Published: (2025)
A Communication-Efficient Decentralized Actor-Critic Algorithm
by: Ren, Xiaoxing, et al.
Published: (2025)
by: Ren, Xiaoxing, et al.
Published: (2025)
Exploiting Similarity for Computation and Communication-Efficient Decentralized Optimization
by: Takezawa, Yuki, et al.
Published: (2025)
by: Takezawa, Yuki, et al.
Published: (2025)
DiLoCoX: A Low-Communication Large-Scale Training Framework for Decentralized Cluster
by: Qi, Ji, et al.
Published: (2025)
by: Qi, Ji, et al.
Published: (2025)
Regularizing Neural Network Training via Identity-wise Discriminative Feature Suppression
by: Chapman, Avraham, et al.
Published: (2022)
by: Chapman, Avraham, et al.
Published: (2022)
Similar Items
-
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025) -
SENTINEL: Stagewise Integrity Verification for Pipeline Parallel Decentralized Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026) -
AsyncMesh: Fully Asynchronous Optimization for Data and Pipeline Parallelism
by: Ajanthan, Thalaiyasingam, et al.
Published: (2026) -
NuMuon: Nuclear-Norm-Constrained Muon for Compressible LLM Training
by: Dolatabadi, Hadi Mohaghegh, et al.
Published: (2026) -
Scaling Prompt Instructed Zero Shot Composed Image Retrieval with Image-Only Data
by: Duan, Yiqun, et al.
Published: (2025)