Understanding Gradient Orthogonalization for Deep Learning via Non-Euclidean Trust-Region Optimization
Fuente:
arXiv
Saved in:
| Main Author: | Kovalev, Dmitry |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Stochastic Non-Smooth Convex Optimization with Unbounded Gradients
by: Kovalev, Dmitry
Published: (2026)
by: Kovalev, Dmitry
Published: (2026)
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025)
by: Kovalev, Dmitry, et al.
Published: (2025)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
Optimal Projection-Free Adaptive SGD for Matrix Optimization
by: Kovalev, Dmitry
Published: (2026)
by: Kovalev, Dmitry
Published: (2026)
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)
by: Su, Weijie
Published: (2025)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
On Linear Convergence in Smooth Convex-Concave Bilinearly-Coupled Saddle-Point Optimization: Lower Bounds and Optimal Algorithms
by: Kovalev, Dmitry, et al.
Published: (2024)
by: Kovalev, Dmitry, et al.
Published: (2024)
Lower Bounds and Optimal Algorithms for Non-Smooth Convex Decentralized Optimization over Time-Varying Networks
by: Kovalev, Dmitry, et al.
Published: (2024)
by: Kovalev, Dmitry, et al.
Published: (2024)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Retraction-Free Decentralized Non-convex Optimization with Orthogonal Constraints
by: Sun, Youbang, et al.
Published: (2024)
by: Sun, Youbang, et al.
Published: (2024)
Muon is Provably Faster with Momentum Variance Reduction
by: Qian, Xun, et al.
Published: (2025)
by: Qian, Xun, et al.
Published: (2025)
Corner Gradient Descent
by: Yarotsky, Dmitry
Published: (2025)
by: Yarotsky, Dmitry
Published: (2025)
A Trust-Region Algorithm for Noisy Equality Constrained Optimization
by: Sun, Shigeng, et al.
Published: (2024)
by: Sun, Shigeng, et al.
Published: (2024)
Non-Euclidean Broximal Point Method: A Blueprint for Geometry-Aware Optimization
by: Gruntkowska, Kaja, et al.
Published: (2025)
by: Gruntkowska, Kaja, et al.
Published: (2025)
Riemannian Optimization for Non-convex Euclidean Distance Geometry with Global Recovery Guarantees
by: Smith, Chandler, et al.
Published: (2024)
by: Smith, Chandler, et al.
Published: (2024)
A Variance-Reduced Stochastic Gradient Tracking Algorithm for Decentralized Optimization with Orthogonality Constraints
by: Wang, Lei, et al.
Published: (2022)
by: Wang, Lei, et al.
Published: (2022)
Accelerated Stochastic ExtraGradient: Mixing Hessian and Gradient Similarity to Reduce Communication in Distributed and Federated Learning
by: Bylinkin, Dmitry, et al.
Published: (2024)
by: Bylinkin, Dmitry, et al.
Published: (2024)
Keep the Momentum: Conservation Laws beyond Euclidean Gradient Flows
by: Marcotte, Sibylle, et al.
Published: (2024)
by: Marcotte, Sibylle, et al.
Published: (2024)
A Non-Monotone Preconditioned Trust-Region Method for Neural Network Training
by: Angino, Andrea, et al.
Published: (2026)
by: Angino, Andrea, et al.
Published: (2026)
LAGO: A Local-Global Optimization Framework Combining Trust Region Methods and Bayesian Optimization
by: Van Dieren, Eliott, et al.
Published: (2026)
by: Van Dieren, Eliott, et al.
Published: (2026)
Adaptive Moment Estimation Optimization Algorithm Using Projection Gradient for Deep Learning
by: Li, Yongqi, et al.
Published: (2025)
by: Li, Yongqi, et al.
Published: (2025)
Adaptive Replication Strategies in Trust-Region-Based Bayesian Optimization of Stochastic Functions
by: Binois, Mickael, et al.
Published: (2025)
by: Binois, Mickael, et al.
Published: (2025)
Geometric Neural Operators (GNPs) for Data-Driven Deep Learning of Non-Euclidean Operators
by: Quackenbush, Blaine, et al.
Published: (2024)
by: Quackenbush, Blaine, et al.
Published: (2024)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
by: Cai, Yuhang, et al.
Published: (2025)
by: Cai, Yuhang, et al.
Published: (2025)
Active Learning of Deep Neural Networks via Gradient-Free Cutting Planes
by: Zhang, Erica, et al.
Published: (2024)
by: Zhang, Erica, et al.
Published: (2024)
An Inexact Weighted Proximal Trust-Region Method
by: Maia, Leandro Farias, et al.
Published: (2026)
by: Maia, Leandro Farias, et al.
Published: (2026)
Non-Euclidean High-Order Smooth Convex Optimization
by: Contreras, Juan Pablo, et al.
Published: (2024)
by: Contreras, Juan Pablo, et al.
Published: (2024)
Enhancing Fractional Gradient Descent with Learned Optimizers
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
Structured Difference-of-Q via Orthogonal Learning
by: Cao, Defu, et al.
Published: (2024)
by: Cao, Defu, et al.
Published: (2024)
Understanding Optimization in Deep Learning with Central Flows
by: Cohen, Jeremy M., et al.
Published: (2024)
by: Cohen, Jeremy M., et al.
Published: (2024)
Stochastic Trust-Region Methods for Over-parameterized Models
by: Yang, Aike, et al.
Published: (2026)
by: Yang, Aike, et al.
Published: (2026)
Local Linear Convergence of Infeasible Optimization with Orthogonal Constraints
by: Sun, Youbang, et al.
Published: (2024)
by: Sun, Youbang, et al.
Published: (2024)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
by: Patitucci, Francisco, et al.
Published: (2026)
by: Patitucci, Francisco, et al.
Published: (2026)
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise
by: Ahn, Kwangjun, et al.
Published: (2024)
by: Ahn, Kwangjun, et al.
Published: (2024)
Fully Stochastic Trust-Region Sequential Quadratic Programming for Equality-Constrained Optimization Problems
by: Fang, Yuchen, et al.
Published: (2022)
by: Fang, Yuchen, et al.
Published: (2022)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
by: Min, Hancheng, et al.
Published: (2025)
by: Min, Hancheng, et al.
Published: (2025)
TRSVR: An Adaptive Stochastic Trust-Region Method with Variance Reduction
by: Fang, Yuchen, et al.
Published: (2026)
by: Fang, Yuchen, et al.
Published: (2026)
More Optimal Fractional-Order Stochastic Gradient Descent for Non-Convex Optimization Problems
by: Partohaghighi, Mohammad, et al.
Published: (2025)
by: Partohaghighi, Mohammad, et al.
Published: (2025)
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
by: Ding, Kuangyu, et al.
Published: (2023)
by: Ding, Kuangyu, et al.
Published: (2023)
Efficient Optimization with Orthogonality Constraint: a Randomized Riemannian Submanifold Method
by: Han, Andi, et al.
Published: (2025)
by: Han, Andi, et al.
Published: (2025)
Similar Items
-
Stochastic Non-Smooth Convex Optimization with Unbounded Gradients
by: Kovalev, Dmitry
Published: (2026) -
Non-Euclidean SGD for Structured Optimization: Unified Analysis and Improved Rates
by: Kovalev, Dmitry, et al.
Published: (2025) -
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025) -
Optimal Projection-Free Adaptive SGD for Matrix Optimization
by: Kovalev, Dmitry
Published: (2026) -
Isotropic Curvature Model for Understanding Deep Learning Optimization: Is Gradient Orthogonalization Optimal?
by: Su, Weijie
Published: (2025)