Clapping: Removing Per-sample Storage for Pipeline Parallel Distributed Optimization with Communication Compression
Fuente:
arXiv
Saved in:
| Main Authors: | Kong, Boao, Huang, Xu, Xu, Yuqi, Liang, Yixuan, Wang, Bin, Yuan, Kun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
by: Kong, Boao, et al.
Published: (2024)
by: Kong, Boao, et al.
Published: (2024)
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
by: Zhu, Shuchen, et al.
Published: (2024)
by: Zhu, Shuchen, et al.
Published: (2024)
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)
by: Xie, Shengping, et al.
Published: (2025)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
BROS: Bias-Corrected Randomized Subspaces for Memory-Efficient Single-Loop Bilevel Optimization
by: Zhang, Hengrui, et al.
Published: (2026)
by: Zhang, Hengrui, et al.
Published: (2026)
Accelerated Distributed Optimization with Compression and Error Feedback
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
GNMR: Runtime Stability Control for Low-Precision Large Language Model Training
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Distributed Bilevel Optimization with Communication Compression
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Subspace Optimization for Efficient Federated Learning under Heterogeneous Data
by: Zhu, Shuchen, et al.
Published: (2026)
by: Zhu, Shuchen, et al.
Published: (2026)
Greedy Low-Rank Gradient Compression for Distributed Learning with Convergence Guarantees
by: Chen, Chuyan, et al.
Published: (2025)
by: Chen, Chuyan, et al.
Published: (2025)
Robust Out-of-Distribution Stochastic Optimization
by: Li, Xianyu, et al.
Published: (2026)
by: Li, Xianyu, et al.
Published: (2026)
Optimal Complexity in Byzantine-Robust Distributed Stochastic Optimization with Data Heterogeneity
by: Shi, Qiankun, et al.
Published: (2025)
by: Shi, Qiankun, et al.
Published: (2025)
Achieving Linear Speedup with ProxSkip in Distributed Stochastic Optimization
by: Guo, Luyao, et al.
Published: (2023)
by: Guo, Luyao, et al.
Published: (2023)
Towards Faster Decentralized Stochastic Optimization with Communication Compression
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Adaptive Batch Size Schedules for Distributed Training of Language Models with Data and Model Parallelism
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
TAMUNA: Doubly Accelerated Distributed Optimization with Local Training, Compression, and Partial Participation
by: Condat, Laurent, et al.
Published: (2023)
by: Condat, Laurent, et al.
Published: (2023)
Constant Stepsize Q-learning: Distributional Convergence, Bias and Extrapolation
by: Zhang, Yixuan, et al.
Published: (2024)
by: Zhang, Yixuan, et al.
Published: (2024)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
BiCoLoR: Communication-Efficient Optimization with Bidirectional Compression and Local Training
by: Condat, Laurent, et al.
Published: (2026)
by: Condat, Laurent, et al.
Published: (2026)
FedSGM: A Unified Framework for Constraint Aware, Bidirectionally Compressed, Multi-Step Federated Optimization
by: Upadhyay, Antesh, et al.
Published: (2026)
by: Upadhyay, Antesh, et al.
Published: (2026)
Convergence of Spectral Descent for Non-smooth Optimization
by: Yang, Yixuan, et al.
Published: (2026)
by: Yang, Yixuan, et al.
Published: (2026)
Natural Hypergradient Descent: Algorithm Design, Convergence Analysis, and Parallel Implementation
by: Kong, Deyi, et al.
Published: (2026)
by: Kong, Deyi, et al.
Published: (2026)
ADMM Algorithms for Residual Network Training: Convergence Analysis and Parallel Implementation
by: Xu, Jintao, et al.
Published: (2023)
by: Xu, Jintao, et al.
Published: (2023)
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
by: Lau, Tim Tsz-Kit, et al.
Published: (2024)
Distributionally Robust Multi-Objective Optimization
by: Yang, Yufeng, et al.
Published: (2026)
by: Yang, Yufeng, et al.
Published: (2026)
Quantum Learning and Estimation for Coordinated Operation between Distribution Networks and Energy Communities
by: Zhuang, Yingrui, et al.
Published: (2025)
by: Zhuang, Yingrui, et al.
Published: (2025)
Online Non-convex Optimization with Long-term Non-convex Constraints
by: Pan, Shijie, et al.
Published: (2023)
by: Pan, Shijie, et al.
Published: (2023)
Conformal Uncertainty Quantification of Electricity Price Predictions for Risk-Averse Storage Arbitrage
by: Alghumayjan, Saud, et al.
Published: (2024)
by: Alghumayjan, Saud, et al.
Published: (2024)
Subspace Optimization for Large Language Models with Convergence Guarantees
by: He, Yutong, et al.
Published: (2024)
by: He, Yutong, et al.
Published: (2024)
Accelerated Methods with Compressed Communications for Distributed Optimization Problems under Data Similarity
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Improving the Worst-Case Bidirectional Communication Complexity for Nonconvex Distributed Optimization under Function Similarity
by: Gruntkowska, Kaja, et al.
Published: (2024)
by: Gruntkowska, Kaja, et al.
Published: (2024)
Robust and Fast Training via Per-Sample Clipping
by: Nobile, Davide, et al.
Published: (2026)
by: Nobile, Davide, et al.
Published: (2026)
Inexact Moreau Envelope Lagrangian Method for Non-Convex Constrained Optimization under Local Error Bound Conditions on Constraint Functions
by: Huang, Yankun, et al.
Published: (2025)
by: Huang, Yankun, et al.
Published: (2025)
Optimization over Sparse Support-Preserving Sets: Two-Step Projection with Global Optimality Guarantees
by: de Vazelhes, William, et al.
Published: (2025)
by: de Vazelhes, William, et al.
Published: (2025)
Distributionally Robust Optimization
by: Kuhn, Daniel, et al.
Published: (2024)
by: Kuhn, Daniel, et al.
Published: (2024)
Momentum Benefits Non-IID Federated Learning Simply and Provably
by: Cheng, Ziheng, et al.
Published: (2023)
by: Cheng, Ziheng, et al.
Published: (2023)
Gradient Flow Sampler-based Distributionally Robust Optimization
by: Xu, Zusen, et al.
Published: (2025)
by: Xu, Zusen, et al.
Published: (2025)
Similar Items
-
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026) -
Decentralized Bilevel Optimization: A Perspective from Transient Iteration Complexity
by: Kong, Boao, et al.
Published: (2024) -
SPARKLE: A Unified Single-Loop Primal-Dual Framework for Decentralized Bilevel Optimization
by: Zhu, Shuchen, et al.
Published: (2024) -
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
by: He, Yutong, et al.
Published: (2023) -
From PowerSGD to PowerSGD+: Low-Rank Gradient Compression for Distributed Optimization with Convergence Guarantees
by: Xie, Shengping, et al.
Published: (2025)