Distributed Low-Communication Training with Decoupled Momentum Optimization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nedelkoski, Sasho, Acker, Alexander, Kao, Odej, Becker, Soeren, Scheinert, Dominik |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
What happens when nanochat meets DiLoCo?
von: Acker, Alexander, et al.
Veröffentlicht: (2025)
von: Acker, Alexander, et al.
Veröffentlicht: (2025)
Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study
von: Wiesner, Philipp, et al.
Veröffentlicht: (2026)
von: Wiesner, Philipp, et al.
Veröffentlicht: (2026)
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
von: Scheinert, Dominik, et al.
Veröffentlicht: (2026)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2026)
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
von: Geldenhuys, Morgan, et al.
Veröffentlicht: (2024)
von: Geldenhuys, Morgan, et al.
Veröffentlicht: (2024)
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)
Daedalus: Self-Adaptive Horizontal Autoscaling for Resource Efficiency of Distributed Stream Processing Systems
von: Pfister, Benjamin J. J., et al.
Veröffentlicht: (2024)
von: Pfister, Benjamin J. J., et al.
Veröffentlicht: (2024)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
von: Coquelin, Daniel, et al.
Veröffentlicht: (2024)
von: Coquelin, Daniel, et al.
Veröffentlicht: (2024)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
von: Yi, Kai, et al.
Veröffentlicht: (2024)
von: Yi, Kai, et al.
Veröffentlicht: (2024)
DeInfoReg: A Decoupled Learning Framework for Better Training Throughput
von: Huang, Zih-Hao, et al.
Veröffentlicht: (2025)
von: Huang, Zih-Hao, et al.
Veröffentlicht: (2025)
Sizey: Memory-Efficient Execution of Scientific Workflow Tasks
von: Bader, Jonathan, et al.
Veröffentlicht: (2024)
von: Bader, Jonathan, et al.
Veröffentlicht: (2024)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
von: Lu, Yunchi, et al.
Veröffentlicht: (2025)
von: Lu, Yunchi, et al.
Veröffentlicht: (2025)
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
von: Dong, Jianbo, et al.
Veröffentlicht: (2024)
von: Dong, Jianbo, et al.
Veröffentlicht: (2024)
Communication Efficient Distributed Training with Distributed Lion
von: Liu, Bo, et al.
Veröffentlicht: (2024)
von: Liu, Bo, et al.
Veröffentlicht: (2024)
Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
von: Choukroun, Yoni, et al.
Veröffentlicht: (2024)
von: Choukroun, Yoni, et al.
Veröffentlicht: (2024)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
von: Liu, Xinyi, et al.
Veröffentlicht: (2025)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
von: Ockerman, Seth, et al.
Veröffentlicht: (2025)
von: Ockerman, Seth, et al.
Veröffentlicht: (2025)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
von: Li, Dacheng, et al.
Veröffentlicht: (2023)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
von: Kadiyala, Divya Kiran, et al.
Veröffentlicht: (2022)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
von: Yao, Jinghan, et al.
Veröffentlicht: (2024)
FedZero: Leveraging Renewable Excess Energy in Federated Learning
von: Wiesner, Philipp, et al.
Veröffentlicht: (2023)
von: Wiesner, Philipp, et al.
Veröffentlicht: (2023)
Domino: Eliminating Communication in LLM Training via Generic Tensor Slicing and Overlapping
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
von: Wang, Guanhua, et al.
Veröffentlicht: (2024)
ATTENTION2D: Communication Efficient Distributed Self-Attention Mechanism
von: Elango, Venmugil
Veröffentlicht: (2025)
von: Elango, Venmugil
Veröffentlicht: (2025)
FSD-Inference: Fully Serverless Distributed Inference with Scalable Cloud Communication
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
von: Oakley, Joe, et al.
Veröffentlicht: (2024)
Communication-Efficient Personalized Federal Graph Learning via Low-Rank Decomposition
von: Liu, Ruyue, et al.
Veröffentlicht: (2024)
von: Liu, Ruyue, et al.
Veröffentlicht: (2024)
A Fast and Flat Federated Learning Method via Weighted Momentum and Sharpness-Aware Minimization
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Efficient Federated Learning Using Dynamic Update and Adaptive Pruning with Momentum on Shared Server Data
von: Liu, Ji, et al.
Veröffentlicht: (2024)
von: Liu, Ji, et al.
Veröffentlicht: (2024)
SpaFL: Communication-Efficient Federated Learning with Sparse Models and Low computational Overhead
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
von: Kim, Minsu, et al.
Veröffentlicht: (2024)
Communication-free Sampling and 4D Hybrid Parallelism for Scalable Mini-batch GNN Training
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
von: Wei, Cunyang, et al.
Veröffentlicht: (2026)
Edge-Cloud Collaborative Computing on Distributed Intelligence and Model Optimization: A Survey
von: Liu, Jing, et al.
Veröffentlicht: (2025)
von: Liu, Jing, et al.
Veröffentlicht: (2025)
Training Time Prediction for Mixed Precision-based Distributed Training
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
von: Kang, Minchul, et al.
Veröffentlicht: (2026)
Predicting the Performance of Scientific Workflow Tasks for Cluster Resource Management: An Overview of the State of the Art
von: Bader, Jonathan, et al.
Veröffentlicht: (2025)
von: Bader, Jonathan, et al.
Veröffentlicht: (2025)
DisagMoE: Computation-Communication overlapped MoE Training via Disaggregated AF-Pipe Parallelism
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
von: Zeng, Zhichen, et al.
Veröffentlicht: (2026)
Communication Optimization for Distributed Training: Architecture, Advances, and Opportunities
von: Wei, Yunze, et al.
Veröffentlicht: (2024)
von: Wei, Yunze, et al.
Veröffentlicht: (2024)
SFPrompt: Communication-Efficient Split Federated Fine-Tuning for Large Pre-Trained Models over Resource-Limited Devices
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
von: Cao, Linxiao, et al.
Veröffentlicht: (2024)
When Foresight Pruning Meets Zeroth-Order Optimization: Efficient Federated Learning for Low-Memory Devices
von: Zhang, Pengyu, et al.
Veröffentlicht: (2024)
von: Zhang, Pengyu, et al.
Veröffentlicht: (2024)
Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
von: Li, Yang, et al.
Veröffentlicht: (2025)
von: Li, Yang, et al.
Veröffentlicht: (2025)
CommunityAI: Towards Community-based Federated Learning
von: Murturi, Ilir, et al.
Veröffentlicht: (2023)
von: Murturi, Ilir, et al.
Veröffentlicht: (2023)
Byzantine-Robust and Communication-Efficient Distributed Learning via Compressed Momentum Filtering
von: Liu, Changxin, et al.
Veröffentlicht: (2024)
von: Liu, Changxin, et al.
Veröffentlicht: (2024)
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
von: Liu, Guanliang, et al.
Veröffentlicht: (2026)
von: Liu, Guanliang, et al.
Veröffentlicht: (2026)
Optimizing the Deployment of Tiny Transformers on Low-Power MCUs
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
von: Jung, Victor J. B., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
What happens when nanochat meets DiLoCo?
von: Acker, Alexander, et al.
Veröffentlicht: (2025) -
Distributed LLM Pretraining During Renewable Curtailment Windows: A Feasibility Study
von: Wiesner, Philipp, et al.
Veröffentlicht: (2026) -
Beyond Microservices: Testing Web-Scale RCA Methods on GPU-Driven LLM Workloads
von: Scheinert, Dominik, et al.
Veröffentlicht: (2026) -
Demeter: Resource-Efficient Distributed Stream Processing under Dynamic Loads with Multi-Configuration Optimization
von: Geldenhuys, Morgan, et al.
Veröffentlicht: (2024) -
Towards a Peer-to-Peer Data Distribution Layer for Efficient and Collaborative Resource Optimization of Distributed Dataflow Applications
von: Scheinert, Dominik, et al.
Veröffentlicht: (2023)