Adaptive Consensus Gradients Aggregation for Scaled Distributed Training
Fuente:
arXiv
Saved in:
| Main Authors: | Choukroun, Yoni, Azoulay, Shlomi, Kisilev, Pavel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SEAFL: Enhancing Efficiency in Semi-Asynchronous Federated Learning through Adaptive Aggregation and Selective Training
by: Islam, Md Sirajul, et al.
Published: (2025)
by: Islam, Md Sirajul, et al.
Published: (2025)
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
by: Ockerman, Seth, et al.
Published: (2025)
by: Ockerman, Seth, et al.
Published: (2025)
Optimal Transport Aggregation for Distributed Mixture-of-Experts
by: Chamroukhi, Faïcel, et al.
Published: (2023)
by: Chamroukhi, Faïcel, et al.
Published: (2023)
Gradient Correction in Federated Learning with Adaptive Optimization
by: Chen, Evan, et al.
Published: (2025)
by: Chen, Evan, et al.
Published: (2025)
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)
by: Li, Hongyang, et al.
Published: (2024)
EASTER: Embedding Aggregation-based Heterogeneous Models Training in Vertical Federated Learning
by: Wang, Shuo, et al.
Published: (2023)
by: Wang, Shuo, et al.
Published: (2023)
TrainVerify: Equivalence-Based Verification for Distributed LLM Training
by: Lu, Yunchi, et al.
Published: (2025)
by: Lu, Yunchi, et al.
Published: (2025)
FedAA: A Reinforcement Learning Perspective on Adaptive Aggregation for Fair and Robust Federated Learning
by: He, Jialuo, et al.
Published: (2024)
by: He, Jialuo, et al.
Published: (2024)
Distributed Low-Communication Training with Decoupled Momentum Optimization
by: Nedelkoski, Sasho, et al.
Published: (2025)
by: Nedelkoski, Sasho, et al.
Published: (2025)
Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs
by: Lin, Jun-Liang, et al.
Published: (2026)
by: Lin, Jun-Liang, et al.
Published: (2026)
AdaptiveLoad: Towards Efficient Video Diffusion Transformer Training
by: Guo, Yucheng, et al.
Published: (2026)
by: Guo, Yucheng, et al.
Published: (2026)
Galvatron: An Automatic Distributed System for Efficient Foundation Model Training
by: Liu, Xinyi, et al.
Published: (2025)
by: Liu, Xinyi, et al.
Published: (2025)
Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive Drafter
by: Hu, Qinghao, et al.
Published: (2025)
by: Hu, Qinghao, et al.
Published: (2025)
Shard the Gradient, Scale the Model: Serverless Federated Aggregation via Gradient Partitioning
by: Barrak, Amine
Published: (2026)
by: Barrak, Amine
Published: (2026)
RollArt: Scaling Agentic RL Training via Disaggregated Infrastructure
by: Gao, Wei, et al.
Published: (2025)
by: Gao, Wei, et al.
Published: (2025)
DISTFLASHATTN: Distributed Memory-efficient Attention for Long-context LLMs Training
by: Li, Dacheng, et al.
Published: (2023)
by: Li, Dacheng, et al.
Published: (2023)
Guard: Scalable Straggler Detection and Node Health Management for Large-Scale Training
by: Liu, Guanliang, et al.
Published: (2026)
by: Liu, Guanliang, et al.
Published: (2026)
BootSeer: Analyzing and Mitigating Initialization Bottlenecks in Large-Scale LLM Training
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores
by: Mei, Zhiyu, et al.
Published: (2023)
by: Mei, Zhiyu, et al.
Published: (2023)
AB-Training: A Communication-Efficient Approach for Distributed Low-Rank Learning
by: Coquelin, Daniel, et al.
Published: (2024)
by: Coquelin, Daniel, et al.
Published: (2024)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
by: Yi, Kai, et al.
Published: (2024)
by: Yi, Kai, et al.
Published: (2024)
Training Ultra Long Context Language Model with Fully Pipelined Distributed Transformer
by: Yao, Jinghan, et al.
Published: (2024)
by: Yao, Jinghan, et al.
Published: (2024)
ALTO: Adaptive LoRA Tuning and Orchestration for Heterogeneous LoRA Training Workloads
by: Zuo, Jingwei, et al.
Published: (2026)
by: Zuo, Jingwei, et al.
Published: (2026)
COMET: A Comprehensive Cluster Design Methodology for Distributed Deep Learning Training
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
by: Kadiyala, Divya Kiran, et al.
Published: (2022)
ByteScale: Efficient Scaling of LLM Training with a 2048K Context Length on More Than 12,000 GPUs
by: Ge, Hao, et al.
Published: (2025)
by: Ge, Hao, et al.
Published: (2025)
Intelligent Sampling of Extreme-Scale Turbulence Datasets for Accurate and Efficient Spatiotemporal Model Training
by: Brewer, Wesley, et al.
Published: (2025)
by: Brewer, Wesley, et al.
Published: (2025)
FusionLLM: A Decentralized LLM Training System on Geo-distributed GPUs with Adaptive Compression
by: Tang, Zhenheng, et al.
Published: (2024)
by: Tang, Zhenheng, et al.
Published: (2024)
FedAgg: Adaptive Federated Learning with Aggregated Gradients
by: Yuan, Wenhao, et al.
Published: (2023)
by: Yuan, Wenhao, et al.
Published: (2023)
FreeScale: Distributed Training for Sequence Recommendation Models with Minimal Scaling Cost
by: Feng, Chenhao, et al.
Published: (2026)
by: Feng, Chenhao, et al.
Published: (2026)
Distributed Graph Neural Network Inference With Just-In-Time Compilation For Industry-Scale Graphs
by: Wu, Xiabao, et al.
Published: (2025)
by: Wu, Xiabao, et al.
Published: (2025)
Piper: Efficient Large-Scale MoE Training via Resource Modeling and Pipelined Hybrid Parallelism
by: Dash, Sajal, et al.
Published: (2026)
by: Dash, Sajal, et al.
Published: (2026)
FedAH: Aggregated Head for Personalized Federated Learning
by: Zhou, Pengzhan, et al.
Published: (2024)
by: Zhou, Pengzhan, et al.
Published: (2024)
Beyond Aggregation: Guiding Clients in Heterogeneous Federated Learning
by: Wang, Zijian, et al.
Published: (2025)
by: Wang, Zijian, et al.
Published: (2025)
Training Time Prediction for Mixed Precision-based Distributed Training
by: Kang, Minchul, et al.
Published: (2026)
by: Kang, Minchul, et al.
Published: (2026)
EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D Parallelism
by: Chen, Yanxi, et al.
Published: (2023)
by: Chen, Yanxi, et al.
Published: (2023)
FedPBS: Proximal-Balanced Scaling Federated Learning Model for Robust Personalized Training for Non-IID Data
by: AbouNassar, Eman M., et al.
Published: (2026)
by: AbouNassar, Eman M., et al.
Published: (2026)
Enhancing Large-Scale AI Training Efficiency: The C4 Solution for Real-Time Anomaly Detection and Communication Optimization
by: Dong, Jianbo, et al.
Published: (2024)
by: Dong, Jianbo, et al.
Published: (2024)
PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep Learning
by: Wang, Yisu, et al.
Published: (2025)
by: Wang, Yisu, et al.
Published: (2025)
Delayed Random Partial Gradient Averaging for Federated Learning
by: Hu, Xinyi
Published: (2024)
by: Hu, Xinyi
Published: (2024)
MiCRO: Near-Zero Cost Gradient Sparsification for Scaling and Accelerating Distributed DNN Training
by: Yoon, Daegun, et al.
Published: (2023)
by: Yoon, Daegun, et al.
Published: (2023)
Similar Items
-
SEAFL: Enhancing Efficiency in Semi-Asynchronous Federated Learning through Adaptive Aggregation and Selective Training
by: Islam, Md Sirajul, et al.
Published: (2025) -
PGT-I: Scaling Spatiotemporal GNNs with Memory-Efficient Distributed Training
by: Ockerman, Seth, et al.
Published: (2025) -
Optimal Transport Aggregation for Distributed Mixture-of-Experts
by: Chamroukhi, Faïcel, et al.
Published: (2023) -
Gradient Correction in Federated Learning with Adaptive Optimization
by: Chen, Evan, et al.
Published: (2025) -
Trustworthiness of Stochastic Gradient Descent in Distributed Learning
by: Li, Hongyang, et al.
Published: (2024)