MicroAdam: Accurate Adaptive Optimization with Low Space Overhead and Provable Convergence
Fuente:
arXiv
Saved in:
| Main Authors: | Modoranu, Ionut-Vlad, Safaryan, Mher, Malinovsky, Grigory, Kurtic, Eldar, Robert, Thomas, Richtarik, Peter, Alistarh, Dan |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
by: Modoranu, Ionut-Vlad, et al.
Published: (2023)
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024)
by: Robert, Thomas, et al.
Published: (2024)
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
by: Modoranu, Ionut-Vlad, et al.
Published: (2025)
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)
by: Wu, Diyuan, et al.
Published: (2024)
DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root Solvers
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
by: Modoranu, Ionut-Vlad, et al.
Published: (2026)
Unified Scaling Laws for Compressed Representations
by: Panferov, Andrei, et al.
Published: (2025)
by: Panferov, Andrei, et al.
Published: (2025)
LoRDO: Distributed Low-Rank Optimization with Infrequent Communication
by: Jovanović, Andrej, et al.
Published: (2026)
by: Jovanović, Andrej, et al.
Published: (2026)
EvoPress: Accurate Dynamic Model Compression via Evolutionary Search
by: Sieberling, Oliver, et al.
Published: (2024)
by: Sieberling, Oliver, et al.
Published: (2024)
CAGE: Curvature-Aware Gradient Estimation For Accurate Quantization-Aware Training
by: Tabesh, Soroush, et al.
Published: (2025)
by: Tabesh, Soroush, et al.
Published: (2025)
Statistically-Lossless Quantization of Large Language Models
by: Helcig, Michael, et al.
Published: (2026)
by: Helcig, Michael, et al.
Published: (2026)
Mathador-LM: A Dynamic Benchmark for Mathematical Reasoning on Large Language Models
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
Towards Robust Scaling Laws for Optimizers
by: Volkova, Alexandra, et al.
Published: (2026)
by: Volkova, Alexandra, et al.
Published: (2026)
GradSkip: Communication-Accelerated Local Gradient Methods with Better Computational Complexity
by: Maranjyan, Artavazd, et al.
Published: (2022)
by: Maranjyan, Artavazd, et al.
Published: (2022)
First Provable Guarantees for Practical Private FL: Beyond Restrictive Assumptions
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling
by: Dadgarnia, Alireza, et al.
Published: (2026)
by: Dadgarnia, Alireza, et al.
Published: (2026)
Streamlining in the Riemannian Realm: Efficient Riemannian Optimization with Loopless Variance Reduction
by: Demidovich, Yury, et al.
Published: (2024)
by: Demidovich, Yury, et al.
Published: (2024)
On Biased Compression for Distributed Learning
by: Beznosikov, Aleksandr, et al.
Published: (2020)
by: Beznosikov, Aleksandr, et al.
Published: (2020)
A Semi-Convergent Stage-Wise Framework with Provable Global Convergence for Adaptive Total Variation Regularization
by: Luo, Liang, et al.
Published: (2025)
by: Luo, Liang, et al.
Published: (2025)
DarwinLM: Evolutionary Structured Pruning of Large Language Models
by: Tang, Shengkun, et al.
Published: (2025)
by: Tang, Shengkun, et al.
Published: (2025)
"Give Me BF16 or Give Me Death"? Accuracy-Performance Trade-Offs in LLM Quantization
by: Kurtic, Eldar, et al.
Published: (2024)
by: Kurtic, Eldar, et al.
Published: (2024)
DP-MicroAdam: Private and Frugal Algorithm for Training and Fine-tuning
by: Hudişteanu, Mihaela, et al.
Published: (2025)
by: Hudişteanu, Mihaela, et al.
Published: (2025)
Distance-Aware Muon: Adaptive Step Scaling for Normalized Optimization
by: Demidovich, Yury, et al.
Published: (2026)
by: Demidovich, Yury, et al.
Published: (2026)
TAMUNA: Doubly Accelerated Distributed Optimization with Local Training, Compression, and Partial Participation
by: Condat, Laurent, et al.
Published: (2023)
by: Condat, Laurent, et al.
Published: (2023)
Optimizers Qualitatively Alter Solutions And We Should Leverage This
by: Pascanu, Razvan, et al.
Published: (2025)
by: Pascanu, Razvan, et al.
Published: (2025)
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
Scalable Mechanistic Neural Networks for Differential Equations and Machine Learning
by: Chen, Jiale, et al.
Published: (2024)
by: Chen, Jiale, et al.
Published: (2024)
Communication-Efficient Gluon in Federated Learning
by: Qian, Xun, et al.
Published: (2026)
by: Qian, Xun, et al.
Published: (2026)
MAST: Model-Agnostic Sparsified Training
by: Demidovich, Yury, et al.
Published: (2023)
by: Demidovich, Yury, et al.
Published: (2023)
Byzantine Robustness and Partial Participation Can Be Achieved at Once: Just Clip Gradient Differences
by: Malinovsky, Grigory, et al.
Published: (2023)
by: Malinovsky, Grigory, et al.
Published: (2023)
Revisiting Stochastic Proximal Point Methods: Generalized Smoothness and Similarity
by: Tovmasyan, Zhirayr, et al.
Published: (2025)
by: Tovmasyan, Zhirayr, et al.
Published: (2025)
Convergence of Adam for Non-convex Objectives: Relaxed Hyperparameters and Non-ergodic Case
by: He, Meixuan, et al.
Published: (2023)
by: He, Meixuan, et al.
Published: (2023)
Provable Low-Rank Tensor-Train Approximations in the Inverse of Large-Scale Structured Matrices
by: Xiao, Chuanfu, et al.
Published: (2025)
by: Xiao, Chuanfu, et al.
Published: (2025)
An implicit regularized enthalpy Lattice Boltzmann Method for the Stefan problem
by: Luddens, Francky, et al.
Published: (2025)
by: Luddens, Francky, et al.
Published: (2025)
ProxiCBO: A Provably Convergent Consensus-Based Method for Composite Optimization
by: Zhang, Haoyu, et al.
Published: (2026)
by: Zhang, Haoyu, et al.
Published: (2026)
Optimization of Adams-type difference formulas in Hilbert space $W_2^{(2,1)}(0,1)$
by: Shadimetov, Kh. M., et al.
Published: (2023)
by: Shadimetov, Kh. M., et al.
Published: (2023)
Byzantine-Robust and Differentially Private Federated Optimization under Weaker Assumptions
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Convergence Analysis of an Adaptive Nonconforming FEM for Phase-Field Dependent Topology Optimization in Stokes Flow
by: Jin, Bangti, et al.
Published: (2025)
by: Jin, Bangti, et al.
Published: (2025)
A Posteriori Error‐Driven Space‐Time Adaptive FEM for Accurate Avascular Tumor and Drug Interaction Modeling
by: Vivek S. Yadav, et al.
Published: (2026)
by: Vivek S. Yadav, et al.
Published: (2026)
Physics Informed Neural Networks for heat conduction with phase change
by: Madir, Bahae-Eddine, et al.
Published: (2024)
by: Madir, Bahae-Eddine, et al.
Published: (2024)
Similar Items
-
MatryoshkaLoRA: Learning Accurate Hierarchical Low-Rank Representations for LLM Fine-Tuning
by: Modoranu, Ionut-Vlad, et al.
Published: (2026) -
Error Feedback Can Accurately Compress Preconditioners
by: Modoranu, Ionut-Vlad, et al.
Published: (2023) -
LDAdam: Adaptive Optimization from Low-Dimensional Gradient Statistics
by: Robert, Thomas, et al.
Published: (2024) -
FFT-based Dynamic Subspace Selection for Low-Rank Adaptive Optimization of Large Language Models
by: Modoranu, Ionut-Vlad, et al.
Published: (2025) -
The Iterative Optimal Brain Surgeon: Faster Sparse Recovery by Leveraging Second-Order Information
by: Wu, Diyuan, et al.
Published: (2024)