Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Sato, Naoki, Izumi, Koshiro, Iiduka, Hideaki |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
by: Sato, Naoki, et al.
Published: (2023)
by: Sato, Naoki, et al.
Published: (2023)
Explicit and Implicit Graduated Optimization in Deep Neural Networks
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
by: Sato, Naoki, et al.
Published: (2026)
by: Sato, Naoki, et al.
Published: (2026)
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)
by: Nagashima, Shuntaro, et al.
Published: (2026)
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search
by: Tsukada, Yuki, et al.
Published: (2023)
by: Tsukada, Yuki, et al.
Published: (2023)
Convergence Bound and Critical Batch Size of Muon Optimizer
by: Sato, Naoki, et al.
Published: (2025)
by: Sato, Naoki, et al.
Published: (2025)
Muon Converges under Heavy-Tailed Noise: Nonconvex Hölder-Smooth Empirical Risk Minimization
by: Iiduka, Hideaki
Published: (2026)
by: Iiduka, Hideaki
Published: (2026)
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum
by: Kamo, Keisuke, et al.
Published: (2025)
by: Kamo, Keisuke, et al.
Published: (2025)
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates
by: Imaizumi, Kento, et al.
Published: (2024)
by: Imaizumi, Kento, et al.
Published: (2024)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Convergence Analysis of SGD under Expected Smoothness
by: Kawamoto, Yuta, et al.
Published: (2025)
by: Kawamoto, Yuta, et al.
Published: (2025)
Accelerating SGDM via Learning Rate and Batch Size Schedules: A Lyapunov-Based Analysis
by: Kondo, Yuichi, et al.
Published: (2025)
by: Kondo, Yuichi, et al.
Published: (2025)
Convergence of Sharpness-Aware Minimization Algorithms using Increasing Batch Size and Decaying Learning Rate
by: Harada, Hinata, et al.
Published: (2024)
by: Harada, Hinata, et al.
Published: (2024)
Both Asymptotic and Non-Asymptotic Convergence of Quasi-Hyperbolic Momentum using Increasing Batch Size
by: Imaizumi, Kento, et al.
Published: (2025)
by: Imaizumi, Kento, et al.
Published: (2025)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
On the Convergence of Adaptive Gradient Methods for Nonconvex Optimization
by: Zhou, Dongruo, et al.
Published: (2018)
by: Zhou, Dongruo, et al.
Published: (2018)
Shuffling Gradient-Based Methods for Nonconvex-Concave Minimax Optimization
by: Tran-Dinh, Quoc, et al.
Published: (2024)
by: Tran-Dinh, Quoc, et al.
Published: (2024)
Compressed Decentralized Momentum Stochastic Gradient Methods for Nonconvex Optimization
by: Liu, Wei, et al.
Published: (2025)
by: Liu, Wei, et al.
Published: (2025)
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning
by: Ding, Kuangyu, et al.
Published: (2023)
by: Ding, Kuangyu, et al.
Published: (2023)
Gradient-Free Method for Heavily Constrained Nonconvex Optimization
by: Shi, Wanli, et al.
Published: (2024)
by: Shi, Wanli, et al.
Published: (2024)
Comparative Study of Neural Network Methods for Solving Topological Solitons
by: Hashimoto, Koji, et al.
Published: (2024)
by: Hashimoto, Koji, et al.
Published: (2024)
Adaptive Lipschitz-Free Conditional Gradient Methods for Stochastic Composite Nonconvex Optimization
by: Yuan, Ganzhao
Published: (2026)
by: Yuan, Ganzhao
Published: (2026)
SAGRAD: A Program for Neural Network Training with Simulated Annealing and the Conjugate Gradient Method
by: Bernal, Javier, et al.
Published: (2025)
by: Bernal, Javier, et al.
Published: (2025)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
by: Ma, Yuhan, et al.
Published: (2024)
by: Ma, Yuhan, et al.
Published: (2024)
Improving Infinitely Deep Bayesian Neural Networks with Nesterov's Accelerated Gradient Method
by: Yu, Chenxu, et al.
Published: (2026)
by: Yu, Chenxu, et al.
Published: (2026)
Faster Gradient-Free Algorithms for Nonsmooth Nonconvex Stochastic Optimization
by: Chen, Lesi, et al.
Published: (2023)
by: Chen, Lesi, et al.
Published: (2023)
Accelerated Gradient Methods for Sparse Statistical Learning with Nonconvex Penalties
by: Yang, Kai, et al.
Published: (2020)
by: Yang, Kai, et al.
Published: (2020)
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
by: Lin, Tianyi, et al.
Published: (2024)
by: Lin, Tianyi, et al.
Published: (2024)
Adaptive Gradient Regularization: A Faster and Generalizable Optimization Technique for Deep Neural Networks
by: Jiang, Huixiu, et al.
Published: (2024)
by: Jiang, Huixiu, et al.
Published: (2024)
Guaranteed Nonconvex Low-Rank Tensor Estimation via Scaled Gradient Descent
by: Wu, Tong
Published: (2025)
by: Wu, Tong
Published: (2025)
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization
by: Xiao, Nachuan, et al.
Published: (2024)
by: Xiao, Nachuan, et al.
Published: (2024)
Enhanced Adaptive Gradient Algorithms for Nonconvex-PL Minimax Optimization
by: Huang, Feihu, et al.
Published: (2023)
by: Huang, Feihu, et al.
Published: (2023)
Reconstructing Deep Neural Networks: Unleashing the Optimization Potential of Natural Gradient Descent
by: Liu, Weihua, et al.
Published: (2024)
by: Liu, Weihua, et al.
Published: (2024)
Variational Stochastic Gradient Descent for Deep Neural Networks
by: Chen, Haotian, et al.
Published: (2024)
by: Chen, Haotian, et al.
Published: (2024)
Learning from Synthetic Data via Provenance-Based Input Gradient Guidance
by: Nagano, Koshiro, et al.
Published: (2026)
by: Nagano, Koshiro, et al.
Published: (2026)
NeuraLSP: An Efficient and Rigorous Neural Left Singular Subspace Preconditioner for Conjugate Gradient Methods
by: Benanti, Alexander, et al.
Published: (2026)
by: Benanti, Alexander, et al.
Published: (2026)
Zeroth-Order Methods for Stochastic Nonconvex Nonsmooth Composite Optimization
by: Chen, Ziyi, et al.
Published: (2025)
by: Chen, Ziyi, et al.
Published: (2025)
Similar Items
-
Using Stochastic Gradient Descent to Smooth Nonconvex Functions: Analysis of Implicit Graduated Optimization
by: Sato, Naoki, et al.
Published: (2023) -
Explicit and Implicit Graduated Optimization in Deep Neural Networks
by: Sato, Naoki, et al.
Published: (2024) -
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
by: Sato, Naoki, et al.
Published: (2024) -
Lipschitz Multiscale Deep Equilibrium Models: A Theoretically Guaranteed and Accelerated Approach
by: Sato, Naoki, et al.
Published: (2026) -
Improved Convergence Rates of Muon Optimizer for Nonconvex Optimization
by: Nagashima, Shuntaro, et al.
Published: (2026)