Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
Fuente:
arXiv
Saved in:
| Main Authors: | Malinovsky, Grigory, Richtárik, Peter, Horváth, Samuel, Gorbunov, Eduard |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAST: Model-Agnostic Sparsified Training
by: Demidovich, Yury, et al.
Published: (2023)
by: Demidovich, Yury, et al.
Published: (2023)
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
Streamlining in the Riemannian Realm: Efficient Riemannian Optimization with Loopless Variance Reduction
by: Demidovich, Yury, et al.
Published: (2024)
by: Demidovich, Yury, et al.
Published: (2024)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
by: Maranjyan, Artavazd, et al.
Published: (2022)
by: Maranjyan, Artavazd, et al.
Published: (2022)
Smoothed Gradient Clipping and Error Feedback for Decentralized Optimization under Symmetric Heavy-Tailed Noise
by: Yu, Shuhua, et al.
Published: (2023)
by: Yu, Shuhua, et al.
Published: (2023)
Towards a Better Theoretical Understanding of Independent Subnetwork Training
by: Shulgin, Egor, et al.
Published: (2023)
by: Shulgin, Egor, et al.
Published: (2023)
Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024)
by: Panferov, Andrei, et al.
Published: (2024)
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
by: Mahran, Ammar, et al.
Published: (2026)
by: Mahran, Ammar, et al.
Published: (2026)
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
by: Condat, Laurent, et al.
Published: (2024)
by: Condat, Laurent, et al.
Published: (2024)
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
by: Tovmasyan, Zhirayr, et al.
Published: (2026)
by: Tovmasyan, Zhirayr, et al.
Published: (2026)
MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times
by: Maranjyan, Artavazd, et al.
Published: (2024)
by: Maranjyan, Artavazd, et al.
Published: (2024)
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
by: Maziane, Yassine, et al.
Published: (2026)
by: Maziane, Yassine, et al.
Published: (2026)
Do We Need Asynchronous SGD? On the Near-Optimality of Synchronous Methods
by: Begunov, Grigory, et al.
Published: (2026)
by: Begunov, Grigory, et al.
Published: (2026)
ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Distributed Stochastic Momentum Tracking with Local Updates: Achieving Optimal Communication and Iteration Complexities
by: Huang, Kun, et al.
Published: (2025)
by: Huang, Kun, et al.
Published: (2025)
Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis
by: Luo, Ruichen, et al.
Published: (2025)
by: Luo, Ruichen, et al.
Published: (2025)
An Optimistic Gradient Tracking Method for Distributed Minimax Optimization
by: Huang, Yan, et al.
Published: (2025)
by: Huang, Yan, et al.
Published: (2025)
Tailoring Gradient Methods for Differentially-Private Distributed Optimization
by: Wang, Yongqiang, et al.
Published: (2022)
by: Wang, Yongqiang, et al.
Published: (2022)
Load Balancing with Network Latencies via Distributed Gradient Descent
by: Balseiro, Santiago R., et al.
Published: (2025)
by: Balseiro, Santiago R., et al.
Published: (2025)
Distributed Difference of Convex Optimization
by: Khatana, Vivek, et al.
Published: (2024)
by: Khatana, Vivek, et al.
Published: (2024)
Decentralized Gradient-Free Methods for Stochastic Non-Smooth Non-Convex Optimization
by: Lin, Zhenwei, et al.
Published: (2023)
by: Lin, Zhenwei, et al.
Published: (2023)
Smoothed Normalization for Efficient Distributed Private Optimization
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
by: Li, Hongpei, et al.
Published: (2025)
by: Li, Hongpei, et al.
Published: (2025)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025)
by: Huertas, Jorge A., et al.
Published: (2025)
Quantize Once, Train Fast: Allreduce-Compatible Compression with Provable Guarantees
by: Xin, Jihao, et al.
Published: (2023)
by: Xin, Jihao, et al.
Published: (2023)
Achieving Near-Optimal Convergence for Distributed Minimax Optimization with Adaptive Stepsizes
by: Huang, Yan, et al.
Published: (2024)
by: Huang, Yan, et al.
Published: (2024)
CONGO: Compressive Online Gradient Optimization
by: Carleton, Jeremy, et al.
Published: (2024)
by: Carleton, Jeremy, et al.
Published: (2024)
A Privacy Preserving Randomized Gossip Algorithm via Controlled Noise Insertion
by: Hanzely, Filip, et al.
Published: (2019)
by: Hanzely, Filip, et al.
Published: (2019)
An Accelerated Distributed Stochastic Gradient Method with Momentum
by: Huang, Kun, et al.
Published: (2024)
by: Huang, Kun, et al.
Published: (2024)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
Efficient Gradient Methods for Distributed Saddle Problems
by: Luo, Ruichen, et al.
Published: (2026)
by: Luo, Ruichen, et al.
Published: (2026)
A First-Order Algorithm for Decentralised Min-Max Problems
by: Malitsky, Yura, et al.
Published: (2023)
by: Malitsky, Yura, et al.
Published: (2023)
Problem-Parameter-Free Decentralized Nonconvex Stochastic Optimization
by: Li, Jiaxiang, et al.
Published: (2024)
by: Li, Jiaxiang, et al.
Published: (2024)
Temporal Parallelisation of the HJB Equation and Continuous-Time Linear Quadratic Control
by: Särkkä, Simo, et al.
Published: (2022)
by: Särkkä, Simo, et al.
Published: (2022)
Decentralized Distributed Optimization for Saddle Point Problems
by: Rogozin, Alexander, et al.
Published: (2021)
by: Rogozin, Alexander, et al.
Published: (2021)
Decentralized Nonsmooth Nonconvex Optimization with Client Sampling
by: Chen, Xinyan, et al.
Published: (2026)
by: Chen, Xinyan, et al.
Published: (2026)
Similar Items
-
MAST: Model-Agnostic Sparsified Training
by: Demidovich, Yury, et al.
Published: (2023) -
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020) -
Streamlining in the Riemannian Realm: Efficient Riemannian Optimization with Loopless Variance Reduction
by: Demidovich, Yury, et al.
Published: (2024) -
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
by: Maranjyan, Artavazd, et al.
Published: (2022) -
Smoothed Gradient Clipping and Error Feedback for Decentralized Optimization under Symmetric Heavy-Tailed Noise
by: Yu, Shuhua, et al.
Published: (2023)