Improving Infinitely Deep Bayesian Neural Networks with Nesterov's Accelerated Gradient Method
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Chenxu, Fang, Wenqi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024)
by: Xu, Zhenghao, et al.
Published: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022)
by: Liu, Xin, et al.
Published: (2022)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024)
by: Park, Chanwoong, et al.
Published: (2024)
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025)
by: Liu, Jun
Published: (2025)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)
by: Liao, Fangshuo, et al.
Published: (2023)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
by: Yau, Chung-Yiu, et al.
Published: (2026)
by: Yau, Chung-Yiu, et al.
Published: (2026)
Accelerated Policy Gradient: On the Convergence Rates of the Nesterov Momentum for Reinforcement Learning
by: Chen, Yen-Ju, et al.
Published: (2023)
by: Chen, Yen-Ju, et al.
Published: (2023)
Partially Stochastic Infinitely Deep Bayesian Neural Networks
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
by: Calvo-Ordonez, Sergio, et al.
Published: (2024)
Inference of Online Newton Methods with Nesterov's Accelerated Sketching
by: Wang, Haoxuan, et al.
Published: (2026)
by: Wang, Haoxuan, et al.
Published: (2026)
Randomized Subspace Nesterov Accelerated Gradient
by: Omiya, Gaku, et al.
Published: (2026)
by: Omiya, Gaku, et al.
Published: (2026)
Nesterov Finds GRAAL: Optimal and Adaptive Gradient Method for Convex Optimization
by: Borodich, Ekaterina, et al.
Published: (2025)
by: Borodich, Ekaterina, et al.
Published: (2025)
Adaptive Nesterov Accelerated Distributional Deep Hedging for Efficient Volatility Risk Management
by: Zhao, Lei, et al.
Published: (2025)
by: Zhao, Lei, et al.
Published: (2025)
SNOO: Step-K Nesterov Outer Optimizer - The Surprising Effectiveness of Nesterov Momentum Applied to Pseudo-Gradients
by: Kallusky, Dominik, et al.
Published: (2025)
by: Kallusky, Dominik, et al.
Published: (2025)
Nesterov-Accelerated Robust Federated Learning Over Byzantine Adversaries
by: Xu, Lihan, et al.
Published: (2025)
by: Xu, Lihan, et al.
Published: (2025)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
by: Vernon, Sydney, et al.
Published: (2025)
by: Vernon, Sydney, et al.
Published: (2025)
Scaled Conjugate Gradient Method for Nonconvex Optimization in Deep Neural Networks
by: Sato, Naoki, et al.
Published: (2024)
by: Sato, Naoki, et al.
Published: (2024)
Functional Stochastic Gradient MCMC for Bayesian Neural Networks
by: Wu, Mengjing, et al.
Published: (2024)
by: Wu, Mengjing, et al.
Published: (2024)
YuriiFormer: A Suite of Nesterov-Accelerated Transformers
by: Zimin, Aleksandr, et al.
Published: (2026)
by: Zimin, Aleksandr, et al.
Published: (2026)
Towards Scalable Bayesian Optimization via Gradient-Informed Bayesian Neural Networks
by: Makrygiorgos, Georgios, et al.
Published: (2025)
by: Makrygiorgos, Georgios, et al.
Published: (2025)
MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training
by: Li, Jiacheng, et al.
Published: (2026)
by: Li, Jiacheng, et al.
Published: (2026)
Nesterov Method for Asynchronous Pipeline Parallel Optimization
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
by: Ajanthan, Thalaiyasingam, et al.
Published: (2025)
Improved Depth Estimation of Bayesian Neural Networks
by: van Erp, Bart, et al.
Published: (2024)
by: van Erp, Bart, et al.
Published: (2024)
DCP: Learning Accelerator Dataflow for Neural Network via Propagation
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
LION-DG: Layer-Informed Initialization with Deep Gradient Protocols for Accelerated Neural Network Training
by: Kim, Hyunjun
Published: (2026)
by: Kim, Hyunjun
Published: (2026)
Posterior Inference on Shallow Infinitely Wide Bayesian Neural Networks under Weights with Unbounded Variance
by: Loría, Jorge, et al.
Published: (2023)
by: Loría, Jorge, et al.
Published: (2023)
Variational Stochastic Gradient Descent for Deep Neural Networks
by: Chen, Haotian, et al.
Published: (2024)
by: Chen, Haotian, et al.
Published: (2024)
Adaptive Stepsizing for Stochastic Gradient Langevin Dynamics in Bayesian Neural Networks
by: Rajpal, Rajit, et al.
Published: (2025)
by: Rajpal, Rajit, et al.
Published: (2025)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
by: Xie, Xingyu, et al.
Published: (2022)
by: Xie, Xingyu, et al.
Published: (2022)
On the Convergence of Locally Adaptive and Scalable Diffusion-Based Sampling Methods for Deep Bayesian Neural Network Posteriors
by: Rensmeyer, Tim, et al.
Published: (2024)
by: Rensmeyer, Tim, et al.
Published: (2024)
Adaptive Gradient Regularization: A Faster and Generalizable Optimization Technique for Deep Neural Networks
by: Jiang, Huixiu, et al.
Published: (2024)
by: Jiang, Huixiu, et al.
Published: (2024)
AdaDPIGU: Differentially Private SGD with Adaptive Clipping and Importance-Based Gradient Updates for Deep Neural Networks
by: Zhang, Huiqi, et al.
Published: (2025)
by: Zhang, Huiqi, et al.
Published: (2025)
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training
by: Ma, Yuhan, et al.
Published: (2024)
by: Ma, Yuhan, et al.
Published: (2024)
TreeLUT: An Efficient Alternative to Deep Neural Networks for Inference Acceleration Using Gradient Boosted Decision Trees
by: Khataei, Alireza, et al.
Published: (2025)
by: Khataei, Alireza, et al.
Published: (2025)
Flexible Infinite-Width Graph Convolutional Neural Networks
by: Anson, Ben, et al.
Published: (2024)
by: Anson, Ben, et al.
Published: (2024)
Infinite Width Limits of Self Supervised Neural Networks
by: Fleissner, Maximilian, et al.
Published: (2024)
by: Fleissner, Maximilian, et al.
Published: (2024)
Improving Generalization of Deep Neural Networks by Optimum Shifting
by: Zhou, Yuyan, et al.
Published: (2024)
by: Zhou, Yuyan, et al.
Published: (2024)
NeuralGrok: Accelerate Grokking by Neural Gradient Transformation
by: Zhou, Xinyu, et al.
Published: (2025)
by: Zhou, Xinyu, et al.
Published: (2025)
Explaining Deep Neural Networks by Leveraging Intrinsic Methods
by: La Rosa, Biagio
Published: (2024)
by: La Rosa, Biagio
Published: (2024)
Sequential Bayesian Optimal Experimental Design in Infinite Dimensions via Policy Gradient Reinforcement Learning
by: Shen, Kaichen, et al.
Published: (2026)
by: Shen, Kaichen, et al.
Published: (2026)
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods
by: Taheri, Hossein, et al.
Published: (2024)
by: Taheri, Hossein, et al.
Published: (2024)
Similar Items
-
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
by: Xu, Zhenghao, et al.
Published: (2024) -
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
by: Liu, Xin, et al.
Published: (2022) -
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
by: Park, Chanwoong, et al.
Published: (2024) -
A Concise Lyapunov Analysis of Nesterov's Accelerated Gradient Method
by: Liu, Jun
Published: (2025) -
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
by: Liao, Fangshuo, et al.
Published: (2023)