Improving Convergence and Generalization Using Parameter Symmetries
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Bo, Gower, Robert M., Walters, Robin, Yu, Rose |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026)
by: Parshakova, Tetiana, et al.
Published: (2026)
Level Set Teleportation: An Optimization Perspective
by: Mishkin, Aaron, et al.
Published: (2024)
by: Mishkin, Aaron, et al.
Published: (2024)
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
by: Sadiev, Abdurakhmon, et al.
Published: (2025)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Improving Generalization and Convergence by Enhancing Implicit Regularization
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
Convergence for Discrete Parameter Update Schemes
by: Wilson, Paul, et al.
Published: (2025)
by: Wilson, Paul, et al.
Published: (2025)
SANIA: Polyak-type Optimization Framework Leads to Scale Invariant Stochastic Algorithms
by: Abdukhakimov, Farshed, et al.
Published: (2023)
by: Abdukhakimov, Farshed, et al.
Published: (2023)
Symmetry in Neural Network Parameter Spaces
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness
by: Bo, Yanan, et al.
Published: (2026)
by: Bo, Yanan, et al.
Published: (2026)
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
by: Schaipp, Fabian, et al.
Published: (2026)
by: Schaipp, Fabian, et al.
Published: (2026)
Learning Provably Improves the Convergence of Gradient Descent
by: Song, Qingyu, et al.
Published: (2025)
by: Song, Qingyu, et al.
Published: (2025)
Random Sparse Lifts: Construction, Analysis and Convergence of finite sparse networks
by: Robin, David A. R., et al.
Published: (2025)
by: Robin, David A. R., et al.
Published: (2025)
Using Taylor-Approximated Gradients to Improve the Frank-Wolfe Method for Empirical Risk Minimization
by: Xiong, Zikai, et al.
Published: (2022)
by: Xiong, Zikai, et al.
Published: (2022)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Optimistic Online-to-Batch Conversions for Accelerated Convergence and Universality
by: Yan, Yu-Hu, et al.
Published: (2025)
by: Yan, Yu-Hu, et al.
Published: (2025)
GANs as Gradient Flows that Converge
by: Huang, Yu-Jui, et al.
Published: (2022)
by: Huang, Yu-Jui, et al.
Published: (2022)
MGDA Converges under Generalized Smoothness, Provably
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Understanding Mode Connectivity via Parameter Space Symmetry
by: Zhao, Bo, et al.
Published: (2025)
by: Zhao, Bo, et al.
Published: (2025)
Improved Last-Iterate Convergence of Shuffling Gradient Methods for Nonsmooth Convex Optimization
by: Liu, Zijian, et al.
Published: (2025)
by: Liu, Zijian, et al.
Published: (2025)
Memory-Reduced Meta-Learning with Guaranteed Convergence
by: Yang, Honglin, et al.
Published: (2024)
by: Yang, Honglin, et al.
Published: (2024)
Towards Fully Parameter-Free Stochastic Optimization: Grid Search with Self-Bounding Analysis
by: Zhao, Yuheng, et al.
Published: (2026)
by: Zhao, Yuheng, et al.
Published: (2026)
Large Deviation Upper Bounds and Improved MSE Rates of Nonlinear SGD: Heavy-tailed Noise and Power of Symmetry
by: Armacki, Aleksandar, et al.
Published: (2024)
by: Armacki, Aleksandar, et al.
Published: (2024)
Revisiting Convergence: Shuffling Complexity Beyond Lipschitz Smoothness
by: He, Qi, et al.
Published: (2025)
by: He, Qi, et al.
Published: (2025)
Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
A Short and Unified Convergence Analysis of the SAG, SAGA, and IAG Algorithms
by: Zhu, Feng, et al.
Published: (2026)
by: Zhu, Feng, et al.
Published: (2026)
Revisiting Subgradient Method: Complexity and Convergence Beyond Lipschitz Continuity
by: Li, Xiao, et al.
Published: (2023)
by: Li, Xiao, et al.
Published: (2023)
Distributed Random Reshuffling Methods with Improved Convergence
by: Huang, Kun, et al.
Published: (2023)
by: Huang, Kun, et al.
Published: (2023)
A Generalization Result for Convergence in Learning-to-Optimize
by: Sucker, Michael, et al.
Published: (2024)
by: Sucker, Michael, et al.
Published: (2024)
On Convergence of Average-Reward Q-Learning in Weakly Communicating Markov Decision Processes
by: Wan, Yi, et al.
Published: (2024)
by: Wan, Yi, et al.
Published: (2024)
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD
by: Jin, Ruinan, et al.
Published: (2024)
by: Jin, Ruinan, et al.
Published: (2024)
On the Global Convergence of Risk-Averse Natural Policy Gradient Methods with Expected Conditional Risk Measures
by: Yu, Xian, et al.
Published: (2023)
by: Yu, Xian, et al.
Published: (2023)
A Provably Convergent Plug-and-Play Framework for Stochastic Bilevel Optimization
by: Chu, Tianshu, et al.
Published: (2025)
by: Chu, Tianshu, et al.
Published: (2025)
Adam-SHANG: A Convergent Adam-Type Method for Stochastic Smooth Convex Optimization
by: Yu, Yaxin, et al.
Published: (2026)
by: Yu, Yaxin, et al.
Published: (2026)
Parameter-free Algorithms for the Stochastically Extended Adversarial Model
by: Wang, Shuche, et al.
Published: (2025)
by: Wang, Shuche, et al.
Published: (2025)
Convergence Guarantees for RMSProp and Adam in Generalized-smooth Non-convex Optimization with Affine Noise Variance
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
Convergence Rate Analysis of LION
by: Dong, Yiming, et al.
Published: (2024)
by: Dong, Yiming, et al.
Published: (2024)
Convergence of Muon with Newton-Schulz
by: Kim, Gyu Yeol, et al.
Published: (2026)
by: Kim, Gyu Yeol, et al.
Published: (2026)
Reusing Historical Trajectories in Natural Policy Gradient via Importance Sampling: Convergence and Convergence Rate
by: Lin, Yifan, et al.
Published: (2024)
by: Lin, Yifan, et al.
Published: (2024)
Similar Items
-
Directional Smoothness and Gradient Methods: Convergence and Adaptivity
by: Mishkin, Aaron, et al.
Published: (2024) -
Muon Does Not Converge on Convex Lipschitz Functions
by: Parshakova, Tetiana, et al.
Published: (2026) -
Level Set Teleportation: An Optimization Perspective
by: Mishkin, Aaron, et al.
Published: (2024) -
Improved Convergence in Parameter-Agnostic Error Feedback through Momentum
by: Sadiev, Abdurakhmon, et al.
Published: (2025) -
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)