Regularized Adaptive Momentum Dual Averaging with an Efficient Inexact Subproblem Solver for Training Structured Neural Network
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Zih-Syuan, Lee, Ching-pei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Adaptive Momentum and Nonlinear Damping for Neural Network Training
by: Karoni, Aikaterini, et al.
Published: (2026)
by: Karoni, Aikaterini, et al.
Published: (2026)
Accelerating nuclear-norm regularized low-rank matrix optimization through Burer-Monteiro decomposition
by: Lee, Ching-pei, et al.
Published: (2022)
by: Lee, Ching-pei, et al.
Published: (2022)
An Adaptively Inexact Method for Bilevel Learning Using Primal-Dual Style Differentiation
by: Bogensperger, Lea, et al.
Published: (2024)
by: Bogensperger, Lea, et al.
Published: (2024)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026)
by: Choudhury, Sayantan, et al.
Published: (2026)
Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
by: Noel, Molly, et al.
Published: (2025)
by: Noel, Molly, et al.
Published: (2025)
Inexact Column Generation for Bayesian Network Structure Learning via Difference-of-Submodular Optimization
by: Yang, Yiran, et al.
Published: (2025)
by: Yang, Yiran, et al.
Published: (2025)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
by: Tucat, Matteo, et al.
Published: (2024)
by: Tucat, Matteo, et al.
Published: (2024)
Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
by: Chen, Cheng, et al.
Published: (2020)
by: Chen, Cheng, et al.
Published: (2020)
Regularized Q-learning through Robust Averaging
by: Schmitt-Förster, Peter, et al.
Published: (2024)
by: Schmitt-Förster, Peter, et al.
Published: (2024)
Dual Natural Gradient Descent for Scalable Training of Physics-Informed Neural Networks
by: Jnini, Anas, et al.
Published: (2025)
by: Jnini, Anas, et al.
Published: (2025)
Leveraging Continuous Time to Understand Momentum When Training Diagonal Linear Networks
by: Papazov, Hristo, et al.
Published: (2024)
by: Papazov, Hristo, et al.
Published: (2024)
DADA: Dual Averaging with Distance Adaptation
by: Moshtaghifar, Mohammad, et al.
Published: (2025)
by: Moshtaghifar, Mohammad, et al.
Published: (2025)
Inexact subgradient methods for semialgebraic functions
by: Bolte, Jérôme, et al.
Published: (2024)
by: Bolte, Jérôme, et al.
Published: (2024)
Bilevel Learning with Inexact Stochastic Gradients
by: Salehi, Mohammad Sadegh, et al.
Published: (2024)
by: Salehi, Mohammad Sadegh, et al.
Published: (2024)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Revisiting Inexact Fixed-Point Iterations for Min-Max Problems: Stochasticity and Structured Nonconvexity
by: Alacaoglu, Ahmet, et al.
Published: (2024)
by: Alacaoglu, Ahmet, et al.
Published: (2024)
An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture
by: Krishnanunni, C G, et al.
Published: (2022)
by: Krishnanunni, C G, et al.
Published: (2022)
An Inexact Weighted Proximal Trust-Region Method
by: Maia, Leandro Farias, et al.
Published: (2026)
by: Maia, Leandro Farias, et al.
Published: (2026)
Beyond the Ideal: Analyzing the Inexact Muon Update
by: Shulgin, Egor, et al.
Published: (2025)
by: Shulgin, Egor, et al.
Published: (2025)
Regularized Gauss-Newton for Optimizing Overparameterized Neural Networks
by: Adeoye, Adeyemi D., et al.
Published: (2024)
by: Adeoye, Adeyemi D., et al.
Published: (2024)
SGD with Adaptive Preconditioning: Unified Analysis and Momentum Acceleration
by: Kovalev, Dmitry
Published: (2025)
by: Kovalev, Dmitry
Published: (2025)
Adaptive Optimization via Momentum on Variance-Normalized Gradients
by: Patitucci, Francisco, et al.
Published: (2026)
by: Patitucci, Francisco, et al.
Published: (2026)
Composite Optimization with Error Feedback: the Dual Averaging Approach
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
Layer-wise Quantization for Quantized Optimistic Dual Averaging
by: Nguyen, Anh Duc, et al.
Published: (2025)
by: Nguyen, Anh Duc, et al.
Published: (2025)
Towards Efficient Constraint Handling in Neural Solvers for Routing Problems
by: Bi, Jieyi, et al.
Published: (2026)
by: Bi, Jieyi, et al.
Published: (2026)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
by: Beneventano, Pierfrancesco, et al.
Published: (2024)
On the Error-Propagation of Inexact Hotelling's Deflation for Principal Component Analysis
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
Unified Convergence Analysis for Adaptive Optimization with Moving Average Estimator
by: Guo, Zhishuai, et al.
Published: (2021)
by: Guo, Zhishuai, et al.
Published: (2021)
Convergence and Complexity Guarantee for Inexact First-order Riemannian Optimization Algorithms
by: Li, Yuchen, et al.
Published: (2024)
by: Li, Yuchen, et al.
Published: (2024)
From Inexact Gradients to Byzantine Robustness: Acceleration and Optimization under Similarity
by: Gaucher, Renaud, et al.
Published: (2026)
by: Gaucher, Renaud, et al.
Published: (2026)
Inexact Moreau Envelope Lagrangian Method for Non-Convex Constrained Optimization under Local Error Bound Conditions on Constraint Functions
by: Huang, Yankun, et al.
Published: (2025)
by: Huang, Yankun, et al.
Published: (2025)
Revisiting Superlinear Convergence of Proximal Newton-Like Methods to Degenerate Solutions
by: Lee, Ching-pei, et al.
Published: (2026)
by: Lee, Ching-pei, et al.
Published: (2026)
Accelerated projected gradient algorithms for sparsity constrained optimization problems
by: Alcantara, Jan Harold, et al.
Published: (2022)
by: Alcantara, Jan Harold, et al.
Published: (2022)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Relaxation-Informed Training of Neural Network Surrogate Models
by: Tsay, Calvin
Published: (2026)
by: Tsay, Calvin
Published: (2026)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
Adaptive Regularized Newton Method with Inexact Hessian
by: Shestakov, Aleksandr, et al.
Published: (2025)
by: Shestakov, Aleksandr, et al.
Published: (2025)
Compression-aware Training of Neural Networks using Frank-Wolfe
by: Zimmer, Max, et al.
Published: (2022)
by: Zimmer, Max, et al.
Published: (2022)
Towards Quantifying the Hessian Structure of Neural Networks
by: Dong, Zhaorui, et al.
Published: (2025)
by: Dong, Zhaorui, et al.
Published: (2025)
Similar Items
-
Adaptive Momentum and Nonlinear Damping for Neural Network Training
by: Karoni, Aikaterini, et al.
Published: (2026) -
Accelerating nuclear-norm regularized low-rank matrix optimization through Burer-Monteiro decomposition
by: Lee, Ching-pei, et al.
Published: (2022) -
An Adaptively Inexact Method for Bilevel Learning Using Primal-Dual Style Differentiation
by: Bogensperger, Lea, et al.
Published: (2024) -
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
by: Choudhury, Sayantan, et al.
Published: (2026) -
Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
by: Noel, Molly, et al.
Published: (2025)