MAST: Model-Agnostic Sparsified Training
Fuente:
arXiv
Saved in:
| Main Authors: | Demidovich, Yury, Malinovsky, Grigory, Shulgin, Egor, Richtárik, Peter |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Streamlining in the Riemannian Realm: Efficient Riemannian Optimization with Loopless Variance Reduction
by: Demidovich, Yury, et al.
Published: (2024)
by: Demidovich, Yury, et al.
Published: (2024)
Towards a Better Theoretical Understanding of Independent Subnetwork Training
by: Shulgin, Egor, et al.
Published: (2023)
by: Shulgin, Egor, et al.
Published: (2023)
Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
by: Malinovsky, Grigory, et al.
Published: (2023)
by: Malinovsky, Grigory, et al.
Published: (2023)
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024)
by: Panferov, Andrei, et al.
Published: (2024)
Smoothed Normalization for Efficient Distributed Private Optimization
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
LoCoDL: Communication-Efficient Distributed Learning with Local Training and Compression
by: Condat, Laurent, et al.
Published: (2024)
by: Condat, Laurent, et al.
Published: (2024)
Ringleader ASGD: The First Asynchronous SGD with Optimal Time Complexity under Data Heterogeneity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
by: Maranjyan, Artavazd, et al.
Published: (2022)
by: Maranjyan, Artavazd, et al.
Published: (2022)
Rescaled Asynchronous SGD: Optimal Distributed Optimization under Data and System Heterogeneity
by: Mahran, Ammar, et al.
Published: (2026)
by: Mahran, Ammar, et al.
Published: (2026)
Rennala MVR: Improved Time Complexity for Parallel Stochastic Optimization via Momentum-Based Variance Reduction
by: Tovmasyan, Zhirayr, et al.
Published: (2026)
by: Tovmasyan, Zhirayr, et al.
Published: (2026)
LOSCAR-SGD: Local SGD with Communication-Computation Overlap and Delay-Corrected Sparse Model Averaging
by: Maziane, Yassine, et al.
Published: (2026)
by: Maziane, Yassine, et al.
Published: (2026)
MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times
by: Maranjyan, Artavazd, et al.
Published: (2024)
by: Maranjyan, Artavazd, et al.
Published: (2024)
Ringmaster LMO: Asynchronous Linear Minimization Oracle Momentum Method
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
by: Sadiev, Abdurakhmon, et al.
Published: (2026)
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
ATA: Adaptive Task Allocation for Efficient Resource Management in Distributed Machine Learning
by: Maranjyan, Artavazd, et al.
Published: (2025)
by: Maranjyan, Artavazd, et al.
Published: (2025)
Communication Efficient Distributed Training with Distributed Lion
by: Liu, Bo, et al.
Published: (2024)
by: Liu, Bo, et al.
Published: (2024)
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
by: Ao, Ruicheng, et al.
Published: (2025)
by: Ao, Ruicheng, et al.
Published: (2025)
FADAS: Towards Federated Adaptive Asynchronous Optimization
by: Wang, Yujia, et al.
Published: (2024)
by: Wang, Yujia, et al.
Published: (2024)
A Communication and Computation Efficient Fully First-order Method for Decentralized Bilevel Optimization
by: Wen, Min, et al.
Published: (2024)
by: Wen, Min, et al.
Published: (2024)
FedComLoc: Communication-Efficient Distributed Training of Sparse and Quantized Models
by: Yi, Kai, et al.
Published: (2024)
by: Yi, Kai, et al.
Published: (2024)
Activations and Gradients Compression for Model-Parallel Training
by: Rudakov, Mikhail, et al.
Published: (2024)
by: Rudakov, Mikhail, et al.
Published: (2024)
Do We Need Asynchronous SGD? On the Near-Optimality of Synchronous Methods
by: Begunov, Grigory, et al.
Published: (2026)
by: Begunov, Grigory, et al.
Published: (2026)
GRAWA: Gradient-based Weighted Averaging for Distributed Training of Deep Learning Models
by: Dimlioglu, Tolga, et al.
Published: (2024)
by: Dimlioglu, Tolga, et al.
Published: (2024)
A Survey of Optimization Methods for Training DL Models: Theoretical Perspective on Convergence and Generalization
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
OptPipe: Memory- and Scheduling-Optimized Pipeline Parallelism for LLM Training
by: Li, Hongpei, et al.
Published: (2025)
by: Li, Hongpei, et al.
Published: (2025)
A Privacy Preserving Randomized Gossip Algorithm via Controlled Noise Insertion
by: Hanzely, Filip, et al.
Published: (2019)
by: Hanzely, Filip, et al.
Published: (2019)
Constraint Programming Models For Serial Batch Scheduling With Minimum Batch Size
by: Huertas, Jorge A., et al.
Published: (2025)
by: Huertas, Jorge A., et al.
Published: (2025)
First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
by: Liao, Fangshuo, et al.
Published: (2025)
by: Liao, Fangshuo, et al.
Published: (2025)
FIARSE: Model-Heterogeneous Federated Learning via Importance-Aware Submodel Extraction
by: Wu, Feijie, et al.
Published: (2024)
by: Wu, Feijie, et al.
Published: (2024)
Stochastic Controlled Averaging for Federated Learning with Communication Compression
by: Huang, Xinmeng, et al.
Published: (2023)
by: Huang, Xinmeng, et al.
Published: (2023)
High-Performance Hybrid Algorithm for Minimum Sum-of-Squares Clustering of Infinitely Tall Data
by: Mussabayev, Ravil, et al.
Published: (2023)
by: Mussabayev, Ravil, et al.
Published: (2023)
Convergence of Sign-based Random Reshuffling Algorithms for Nonconvex Optimization
by: Qin, Zhen, et al.
Published: (2023)
by: Qin, Zhen, et al.
Published: (2023)
Communication-Efficient Federated Bilevel Optimization with Local and Global Lower Level Problems
by: Li, Junyi, et al.
Published: (2023)
by: Li, Junyi, et al.
Published: (2023)
Dynamic Regularized Sharpness Aware Minimization in Federated Learning: Approaching Global Consistency and Smooth Landscape
by: Sun, Yan, et al.
Published: (2023)
by: Sun, Yan, et al.
Published: (2023)
Online Distributed Learning with Quantized Finite-Time Coordination
by: Bastianello, Nicola, et al.
Published: (2023)
by: Bastianello, Nicola, et al.
Published: (2023)
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix
by: Yue, Yun, et al.
Published: (2023)
by: Yue, Yun, et al.
Published: (2023)
Lower Bounds and Accelerated Algorithms in Distributed Stochastic Optimization with Communication Compression
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
Unbiased Compression Saves Communication in Distributed Optimization: When and How Much?
by: He, Yutong, et al.
Published: (2023)
by: He, Yutong, et al.
Published: (2023)
Similar Items
-
Streamlining in the Riemannian Realm: Efficient Riemannian Optimization with Loopless Variance Reduction
by: Demidovich, Yury, et al.
Published: (2024) -
Towards a Better Theoretical Understanding of Independent Subnetwork Training
by: Shulgin, Egor, et al.
Published: (2023) -
Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
by: Malinovsky, Grigory, et al.
Published: (2023) -
Correlated Quantization for Faster Nonconvex Distributed Optimization
by: Panferov, Andrei, et al.
Published: (2024) -
Smoothed Normalization for Efficient Distributed Private Optimization
by: Shulgin, Egor, et al.
Published: (2025)