Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
Fuente:
arXiv
Salvato in:
| Autori principali: | Liao, Fangshuo, Kyrillidis, Anastasios |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
di: Liao, Fangshuo, et al.
Pubblicazione: (2025)
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
di: Liao, Fangshuo, et al.
Pubblicazione: (2026)
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
di: Kolomvaki, Afroditi, et al.
Pubblicazione: (2026)
di: Kolomvaki, Afroditi, et al.
Pubblicazione: (2026)
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
di: Vandchali, Mahtab Alizadeh, et al.
Pubblicazione: (2025)
di: Vandchali, Mahtab Alizadeh, et al.
Pubblicazione: (2025)
On the Error-Propagation of Inexact Hotelling's Deflation for Principal Component Analysis
di: Liao, Fangshuo, et al.
Pubblicazione: (2023)
di: Liao, Fangshuo, et al.
Pubblicazione: (2023)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
di: Xu, Zhenghao, et al.
Pubblicazione: (2024)
di: Xu, Zhenghao, et al.
Pubblicazione: (2024)
Convex Formulations for Training Two-Layer ReLU Neural Networks
di: Prakhya, Karthik, et al.
Pubblicazione: (2024)
di: Prakhya, Karthik, et al.
Pubblicazione: (2024)
Hidden Minima in Two-Layer ReLU Networks
di: Arjevani, Yossi
Pubblicazione: (2023)
di: Arjevani, Yossi
Pubblicazione: (2023)
Stability and Performance Analysis of Discrete-Time ReLU Recurrent Neural Networks
di: Noori, Sahel Vahedi, et al.
Pubblicazione: (2024)
di: Noori, Sahel Vahedi, et al.
Pubblicazione: (2024)
Convex Relaxations of ReLU Neural Networks Approximate Global Optima in Polynomial Time
di: Kim, Sungyoon, et al.
Pubblicazione: (2024)
di: Kim, Sungyoon, et al.
Pubblicazione: (2024)
Exploiting Low-Rank Structure in Max-K-Cut Problems
di: Stevens, Ria, et al.
Pubblicazione: (2026)
di: Stevens, Ria, et al.
Pubblicazione: (2026)
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
di: Yang, Yahong, et al.
Pubblicazione: (2023)
di: Yang, Yahong, et al.
Pubblicazione: (2023)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
di: Min, Hancheng, et al.
Pubblicazione: (2025)
di: Min, Hancheng, et al.
Pubblicazione: (2025)
Computational Tradeoffs of Optimization-Based Bound Tightening in ReLU Networks
di: Badilla, Fabian, et al.
Pubblicazione: (2023)
di: Badilla, Fabian, et al.
Pubblicazione: (2023)
When is Momentum Extragradient Optimal? A Polynomial-Based Analysis
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
di: Kim, Junhyung Lyle, et al.
Pubblicazione: (2022)
ReLU Networks for Model Predictive Control: Network Complexity and Performance Guarantees
di: Li, Xingchen, et al.
Pubblicazione: (2026)
di: Li, Xingchen, et al.
Pubblicazione: (2026)
Geometry-induced Regularization in Deep ReLU Neural Networks
di: Bona-Pellissier, Joachim, et al.
Pubblicazione: (2024)
di: Bona-Pellissier, Joachim, et al.
Pubblicazione: (2024)
Why Smooth Stability Assumptions Fail for ReLU Learning
di: Katende, Ronald
Pubblicazione: (2025)
di: Katende, Ronald
Pubblicazione: (2025)
An analysis of optimization problems involving ReLU neural networks
di: Plate, Christoph, et al.
Pubblicazione: (2025)
di: Plate, Christoph, et al.
Pubblicazione: (2025)
EMA-Nesterov: Stabilizing Nesterov's Lookahead for Accelerated Deep Learning Optimization
di: Yau, Chung-Yiu, et al.
Pubblicazione: (2026)
di: Yau, Chung-Yiu, et al.
Pubblicazione: (2026)
Nonlinear Dynamics In Optimization Landscape of Shallow Neural Networks with Tunable Leaky ReLU
di: Liu, Jingzhou
Pubblicazione: (2025)
di: Liu, Jingzhou
Pubblicazione: (2025)
Approximation with Random Shallow ReLU Networks with Applications to Model Reference Adaptive Control
di: Lamperski, Andrew, et al.
Pubblicazione: (2024)
di: Lamperski, Andrew, et al.
Pubblicazione: (2024)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
di: Lai, Kuo-Wei, et al.
Pubblicazione: (2026)
di: Lai, Kuo-Wei, et al.
Pubblicazione: (2026)
An Efficient Alternating Algorithm for ReLU-based Symmetric Matrix Decomposition
di: Wang, Qingsong
Pubblicazione: (2025)
di: Wang, Qingsong
Pubblicazione: (2025)
A Complete Set of Quadratic Constraints for Repeated ReLU and Generalizations
di: Noori, Sahel Vahedi, et al.
Pubblicazione: (2024)
di: Noori, Sahel Vahedi, et al.
Pubblicazione: (2024)
Adversarial Training of Two-Layer Polynomial and ReLU Activation Networks via Convex Optimization
di: Kuelbs, Daniel, et al.
Pubblicazione: (2024)
di: Kuelbs, Daniel, et al.
Pubblicazione: (2024)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
di: Liu, Xin, et al.
Pubblicazione: (2022)
di: Liu, Xin, et al.
Pubblicazione: (2022)
Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models
di: Xie, Xingyu, et al.
Pubblicazione: (2022)
di: Xie, Xingyu, et al.
Pubblicazione: (2022)
MIQCQP reformulation of the ReLU neural networks Lipschitz constant estimation problem
di: Sbihi, Mohammed, et al.
Pubblicazione: (2024)
di: Sbihi, Mohammed, et al.
Pubblicazione: (2024)
Function Gradient Approximation with Random Shallow ReLU Networks with Control Applications
di: Lamperski, Andrew, et al.
Pubblicazione: (2024)
di: Lamperski, Andrew, et al.
Pubblicazione: (2024)
Path-conditioned training: a principled way to rescale ReLU neural networks
di: Lebeurrier, Arthur, et al.
Pubblicazione: (2026)
di: Lebeurrier, Arthur, et al.
Pubblicazione: (2026)
Local Lipschitz Constant Computation of ReLU-FNNs: Upper Bound Computation with Exactness Verification
di: Ebihara, Yoshio, et al.
Pubblicazione: (2023)
di: Ebihara, Yoshio, et al.
Pubblicazione: (2023)
One Model, Two Roles: Emergent Specialization in a Shared Recurrent Transformer
di: Shen, Jucheng, et al.
Pubblicazione: (2026)
di: Shen, Jucheng, et al.
Pubblicazione: (2026)
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
di: Choudhury, Sayantan, et al.
Pubblicazione: (2026)
di: Choudhury, Sayantan, et al.
Pubblicazione: (2026)
Unveiling Hidden Pivotal Players with GoalNet: A GNN-Based Soccer Player Evaluation System
di: Jiang, Jacky Hao, et al.
Pubblicazione: (2025)
di: Jiang, Jacky Hao, et al.
Pubblicazione: (2025)
Constructive Universal Approximation and Finite Sample Memorization by Narrow Deep ReLU Networks
di: Hernández, Martín, et al.
Pubblicazione: (2024)
di: Hernández, Martín, et al.
Pubblicazione: (2024)
Nesterov Acceleration for Ensemble Kalman Inversion and Variants
di: Vernon, Sydney, et al.
Pubblicazione: (2025)
di: Vernon, Sydney, et al.
Pubblicazione: (2025)
Generalized Continuous-Time Models for Nesterov's Accelerated Gradient Methods
di: Park, Chanwoong, et al.
Pubblicazione: (2024)
di: Park, Chanwoong, et al.
Pubblicazione: (2024)
On bounds for norms of reparameterized ReLU artificial neural network parameters: sums of fractional powers of the Lipschitz norm control the network parameter vector
di: Jentzen, Arnulf, et al.
Pubblicazione: (2022)
di: Jentzen, Arnulf, et al.
Pubblicazione: (2022)
Documenti analoghi
-
Guided by the Experts: Provable Feature Learning Dynamic of Soft-Routed Mixture-of-Experts
di: Liao, Fangshuo, et al.
Pubblicazione: (2025) -
Provable Model-Parallel Distributed Principal Component Analysis with Parallel Deflation
di: Liao, Fangshuo, et al.
Pubblicazione: (2025) -
SGD at the Edge of Stability: The Stochastic Sharpness Gap
di: Liao, Fangshuo, et al.
Pubblicazione: (2026) -
Convergence Analysis of Two-Layer Neural Networks under Gaussian Input Masking
di: Kolomvaki, Afroditi, et al.
Pubblicazione: (2026) -
One Rank at a Time: Cascading Error Dynamics in Sequential Learning
di: Vandchali, Mahtab Alizadeh, et al.
Pubblicazione: (2025)