ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM Training
Fuente:
arXiv
Saved in:
| Main Authors: | Nabli, Adel, Fournier, Louis, Erbacher, Pierre, Serrano, Louis, Belilovsky, Eugene, Oyallon, Edouard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
by: Fournier, Louis, et al.
Published: (2024)
by: Fournier, Louis, et al.
Published: (2024)
Cyclic Data Parallelism for Efficient Parallelism of Deep Neural Networks
by: Fournier, Louis, et al.
Published: (2024)
by: Fournier, Louis, et al.
Published: (2024)
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
PETRA: Parallel End-to-end Training with Reversible Architectures
by: Rivaud, Stéphane, et al.
Published: (2024)
by: Rivaud, Stéphane, et al.
Published: (2024)
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)
by: Janson, Paul, et al.
Published: (2026)
DISCO: learning to DISCover an evolution Operator for multi-physics-agnostic prediction
by: Morel, Rudy, et al.
Published: (2025)
by: Morel, Rudy, et al.
Published: (2025)
Efficient Refusal Ablation in LLM through Optimal Transport
by: Nanfack, Geraldin, et al.
Published: (2026)
by: Nanfack, Geraldin, et al.
Published: (2026)
Warming Up for Zeroth-Order Federated Pre-Training with Low Resource Clients
by: Legate, Gwen, et al.
Published: (2025)
by: Legate, Gwen, et al.
Published: (2025)
Celo2: Towards Learned Optimization Free Lunch
by: Moudgil, Abhinav, et al.
Published: (2026)
by: Moudgil, Abhinav, et al.
Published: (2026)
Accelerating Training with Neuron Interaction and Nowcasting Networks
by: Knyazev, Boris, et al.
Published: (2024)
by: Knyazev, Boris, et al.
Published: (2024)
Not Only the Last-Layer Features for Spurious Correlations: All Layer Deep Feature Reweighting
by: Hameed, Humza Wajid, et al.
Published: (2024)
by: Hameed, Humza Wajid, et al.
Published: (2024)
DiffuMamba: High-Throughput Diffusion LMs with Mamba Backbone
by: Singh, Vaibhav, et al.
Published: (2025)
by: Singh, Vaibhav, et al.
Published: (2025)
On Non-Linear operators for Geometric Deep Learning
by: Sergeant-Perthuis, Grégoire, et al.
Published: (2022)
by: Sergeant-Perthuis, Grégoire, et al.
Published: (2022)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
by: Wang, Guanhua, et al.
Published: (2024)
by: Wang, Guanhua, et al.
Published: (2024)
Less is More: Undertraining Experts Improves Model Upcycling
by: Horoi, Stefan, et al.
Published: (2025)
by: Horoi, Stefan, et al.
Published: (2025)
$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers
by: Thérien, Benjamin, et al.
Published: (2024)
by: Thérien, Benjamin, et al.
Published: (2024)
Test-time Generalization for Physics through Neural Operator Splitting
by: Serrano, Louis, et al.
Published: (2026)
by: Serrano, Louis, et al.
Published: (2026)
Non-Uniform Parameter-Wise Model Merging
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
by: Camacho, Albert Manuel Orozco, et al.
Published: (2024)
Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL
by: Miahi, Erfan, et al.
Published: (2026)
by: Miahi, Erfan, et al.
Published: (2026)
Communication Efficient LLM Pre-training with SparseLoCo
by: Sarfi, Amir, et al.
Published: (2025)
by: Sarfi, Amir, et al.
Published: (2025)
PC Agent: While You Sleep, AI Works -- A Cognitive Journey into Digital World
by: He, Yanheng, et al.
Published: (2024)
by: He, Yanheng, et al.
Published: (2024)
GreedySnake: Accelerating SSD-Offloaded LLM Training with Efficient Scheduling and Optimizer Step Overlapping
by: Yin, Yishu, et al.
Published: (2025)
by: Yin, Yishu, et al.
Published: (2025)
Harmony in Diversity: Merging Neural Networks with Canonical Correlation Analysis
by: Horoi, Stefan, et al.
Published: (2024)
by: Horoi, Stefan, et al.
Published: (2024)
NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
by: Bi, Xiaohan, et al.
Published: (2025)
by: Bi, Xiaohan, et al.
Published: (2025)
Whatever Remains Must Be True: Filtering Drives Reasoning in LLMs, Shaping Diversity
by: Kruszewski, Germán, et al.
Published: (2025)
by: Kruszewski, Germán, et al.
Published: (2025)
Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver
by: Marty, Tom, et al.
Published: (2023)
by: Marty, Tom, et al.
Published: (2023)
Understanding the Role of Training Data in Test-Time Scaling
by: Javanmard, Adel, et al.
Published: (2025)
by: Javanmard, Adel, et al.
Published: (2025)
Does Your Optimizer Care How You Normalize? Normalization-Optimizer Coupling in LLM Training
by: Abouzeid, Abdelrahman
Published: (2026)
by: Abouzeid, Abdelrahman
Published: (2026)
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
by: Hu, Shengyuan, et al.
Published: (2025)
by: Hu, Shengyuan, et al.
Published: (2025)
Scalable Federated Unlearning via Isolated and Coded Sharding
by: Lin, Yijing, et al.
Published: (2024)
by: Lin, Yijing, et al.
Published: (2024)
Theoretical Perspectives on Data Quality and Synergistic Effects in Pre- and Post-Training Reasoning Models
by: Javanmard, Adel, et al.
Published: (2026)
by: Javanmard, Adel, et al.
Published: (2026)
Towards Low-bit Communication for Tensor Parallel LLM Inference
by: Dong, Harry, et al.
Published: (2024)
by: Dong, Harry, et al.
Published: (2024)
AB-UPT for Automotive and Aerospace Applications
by: Alkin, Benedikt, et al.
Published: (2025)
by: Alkin, Benedikt, et al.
Published: (2025)
The Geometries of Truth Are Orthogonal Across Tasks
by: Azizian, Waiss, et al.
Published: (2025)
by: Azizian, Waiss, et al.
Published: (2025)
ProgFed: Effective, Communication, and Computation Efficient Federated Learning by Progressive Training
by: Wang, Hui-Po, et al.
Published: (2021)
by: Wang, Hui-Po, et al.
Published: (2021)
Think While You Generate: Discrete Diffusion with Planned Denoising
by: Liu, Sulin, et al.
Published: (2024)
by: Liu, Sulin, et al.
Published: (2024)
PersonaAgent with GraphRAG: Community-Aware Knowledge Graphs for Personalized LLM
by: Liang, Siqi, et al.
Published: (2025)
by: Liang, Siqi, et al.
Published: (2025)
How Fast Should a Model Commit to Supervision? Training Reasoning Models on the Tsallis Loss Continuum
by: Lin, Chu-Cheng, et al.
Published: (2026)
by: Lin, Chu-Cheng, et al.
Published: (2026)
Strategies for Improving Communication Efficiency in Distributed and Federated Learning: Compression, Local Training, and Personalization
by: Yi, Kai
Published: (2025)
by: Yi, Kai
Published: (2025)
Contrastive Language-Image Pre-Training Model based Semantic Communication Performance Optimization
by: Yang, Shaoran, et al.
Published: (2025)
by: Yang, Shaoran, et al.
Published: (2025)
Similar Items
-
WASH: Train your Ensemble with Communication-Efficient Weight Shuffling, then Average
by: Fournier, Louis, et al.
Published: (2024) -
Cyclic Data Parallelism for Efficient Parallelism of Deep Neural Networks
by: Fournier, Louis, et al.
Published: (2024) -
Model Parallelism With Subnetwork Data Parallelism
by: Singh, Vaibhav, et al.
Published: (2025) -
PETRA: Parallel End-to-end Training with Reversible Architectures
by: Rivaud, Stéphane, et al.
Published: (2024) -
Stabilizing Native Low-Rank LLM Pretraining
by: Janson, Paul, et al.
Published: (2026)