MoMo: Momentum Models for Adaptive Learning Rates
Fuente:
arXiv
Saved in:
| Main Authors: | Schaipp, Fabian, Ohana, Ruben, Eickenberg, Michael, Defazio, Aaron, Gower, Robert M. |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation
by: Gower, Robert M., et al.
Published: (2025)
by: Gower, Robert M., et al.
Published: (2025)
Cutting Some Slack for SGD with Adaptive Polyak Stepsizes
by: Gower, Robert M., et al.
Published: (2022)
by: Gower, Robert M., et al.
Published: (2022)
Structured Sketching for Linear Systems
by: Brust, Johannes J, et al.
Published: (2024)
by: Brust, Johannes J, et al.
Published: (2024)
Stable gradient-adjusted root mean square propagation on least squares problem
by: Li, Runze, et al.
Published: (2024)
by: Li, Runze, et al.
Published: (2024)
Stochastic versus Deterministic in Stochastic Gradient Descent
by: Li, Runze, et al.
Published: (2025)
by: Li, Runze, et al.
Published: (2025)
Stochastic trace estimation for parameter-dependent matrices applied to spectral density approximation
by: Matti, Fabio, et al.
Published: (2025)
by: Matti, Fabio, et al.
Published: (2025)
Federated Learning with Convex Global and Local Constraints
by: He, Chuan, et al.
Published: (2023)
by: He, Chuan, et al.
Published: (2023)
Subspace-constrained randomized coordinate descent for linear systems with good low-rank matrix approximations
by: Lok, Jackie, et al.
Published: (2025)
by: Lok, Jackie, et al.
Published: (2025)
Have ASkotch: A Neat Solution for Large-scale Kernel Ridge Regression
by: Rathore, Pratik, et al.
Published: (2024)
by: Rathore, Pratik, et al.
Published: (2024)
On the boundedness of the sequence generated by minibatch stochastic gradient descent
by: Bauschke, Heinz H., et al.
Published: (2025)
by: Bauschke, Heinz H., et al.
Published: (2025)
Dimension-free estimators of gradients of functions with(out) non-independent variables
by: Lamboni, Matieyendou
Published: (2025)
by: Lamboni, Matieyendou
Published: (2025)
Acceleration and restart for the randomized Bregman-Kaczmarz method
by: Tondji, Lionel, et al.
Published: (2023)
by: Tondji, Lionel, et al.
Published: (2023)
Parallelization Strategies for the Randomized Kaczmarz Algorithm on Large-Scale Dense Systems
by: Ferreira, Inês, et al.
Published: (2024)
by: Ferreira, Inês, et al.
Published: (2024)
Primal-Dual Coordinate Descent for Nonconvex-Nonconcave Saddle Point Problems Under the Weak MVI Assumption
by: Walwil, Iyad, et al.
Published: (2025)
by: Walwil, Iyad, et al.
Published: (2025)
Distributed Computing for Huge-Scale Aggregative Convex Programming
by: Tao, Luoyi
Published: (2026)
by: Tao, Luoyi
Published: (2026)
Classification by Separating Hypersurfaces: An Entropic Approach
by: Arratia, Argimiro, et al.
Published: (2025)
by: Arratia, Argimiro, et al.
Published: (2025)
A Fast Monte Carlo algorithm for evaluating matrix functions with application in complex networks
by: Guidotti, Nicolas L., et al.
Published: (2023)
by: Guidotti, Nicolas L., et al.
Published: (2023)
Some Remarks on the Optimal Level of Randomization in Global Optimization
by: Theodosopoulos, Ted
Published: (2004)
by: Theodosopoulos, Ted
Published: (2004)
Enhancing Diversity in Multi-objective Feature Selection
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
by: Miyandoab, Sevil Zanjani, et al.
Published: (2024)
Adaptive monotonicity testing in sublinear time
by: Li, Housen, et al.
Published: (2025)
by: Li, Housen, et al.
Published: (2025)
An accelerated randomized Bregman-Kaczmarz method for strongly convex linearly constraint optimization
by: Tondji, Lionel, et al.
Published: (2025)
by: Tondji, Lionel, et al.
Published: (2025)
CompressedScaffnew: The First Theoretical Double Acceleration of Communication from Local Training and Compression in Distributed Optimization
by: Condat, Laurent, et al.
Published: (2022)
by: Condat, Laurent, et al.
Published: (2022)
Optimal Online Bipartite Matching in Degree-2 Graphs
by: Bhangale, Amey, et al.
Published: (2025)
by: Bhangale, Amey, et al.
Published: (2025)
RA-DCA: A Randomized Active-Set DCA for Directional Stationarity in Max-Structured DC Programs
by: Niu, Yi-Shuai
Published: (2026)
by: Niu, Yi-Shuai
Published: (2026)
General Constrained Matrix Optimization
by: Garner, Casey, et al.
Published: (2024)
by: Garner, Casey, et al.
Published: (2024)
Spectrally Constrained Optimization
by: Garner, Casey, et al.
Published: (2023)
by: Garner, Casey, et al.
Published: (2023)
Exact recovery for seeded graph matching
by: Fraiman, Nicolas, et al.
Published: (2026)
by: Fraiman, Nicolas, et al.
Published: (2026)
NOVAK: Unified adaptive optimizer for deep neural networks
by: Kavun, Sergii
Published: (2026)
by: Kavun, Sergii
Published: (2026)
On the fast convergence of minibatch heavy ball momentum
by: Bollapragada, Raghu, et al.
Published: (2022)
by: Bollapragada, Raghu, et al.
Published: (2022)
HPR-QP: A dual Halpern Peaceman-Rachford method for solving large-scale convex composite quadratic programming
by: Chen, Kaihuang, et al.
Published: (2025)
by: Chen, Kaihuang, et al.
Published: (2025)
A New Algorithm for Computing Integer Hulls of 2D Polyhedral Sets
by: Mukherjee, Chirantan
Published: (2025)
by: Mukherjee, Chirantan
Published: (2025)
Surrogate-based Autotuning for Randomized Sketching Algorithms in Regression Problems
by: Cho, Younghyun, et al.
Published: (2023)
by: Cho, Younghyun, et al.
Published: (2023)
A Heuristic Alternating Direction Method of Multipliers Framework for Distributed and Centralized Tree-Constrained Optimization: Applications to Hop-Constrained Spanning Tree Multicommodity Flow Design
by: Mokhtari, Yacine
Published: (2025)
by: Mokhtari, Yacine
Published: (2025)
Efficient Multi-Processor Scheduling in Increasingly Realistic Models
by: Papp, Pál András, et al.
Published: (2024)
by: Papp, Pál András, et al.
Published: (2024)
Nature-Inspired Algorithms in Optimization: Introduction, Hybridization and Insights
by: Yang, Xin-She
Published: (2023)
by: Yang, Xin-She
Published: (2023)
Deterministic computation of quantiles in a Lipschitz framework
by: Gu, Yurun, et al.
Published: (2024)
by: Gu, Yurun, et al.
Published: (2024)
Correcting the Foundational Analysis of Karp--Vazirani--Vazirani (STOC 1990): A Rigorous Revision of the $1-1/e$ Upper Bound
by: Xu, Pan
Published: (2025)
by: Xu, Pan
Published: (2025)
Weighted domination models and randomized heuristics
by: Dijkstra, Lukas, et al.
Published: (2022)
by: Dijkstra, Lukas, et al.
Published: (2022)
Adaptive first-order methods with enhanced worst-case rates
by: Florea, Mihai I.
Published: (2024)
by: Florea, Mihai I.
Published: (2024)
Floating Point Compression of Hierarchical Matrix Formats and its Impact on Matrix-Vector Multiplication
by: Kriemann, Ronald
Published: (2024)
by: Kriemann, Ronald
Published: (2024)
Similar Items
-
Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation
by: Gower, Robert M., et al.
Published: (2025) -
Cutting Some Slack for SGD with Adaptive Polyak Stepsizes
by: Gower, Robert M., et al.
Published: (2022) -
Structured Sketching for Linear Systems
by: Brust, Johannes J, et al.
Published: (2024) -
Stable gradient-adjusted root mean square propagation on least squares problem
by: Li, Runze, et al.
Published: (2024) -
Stochastic versus Deterministic in Stochastic Gradient Descent
by: Li, Runze, et al.
Published: (2025)