Gradient descent provably escapes saddle points in the training of shallow ReLU networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cheridito, Patrick, Jentzen, Arnulf, Rossmannek, Florian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2022
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
von: Do, Thang, et al.
Veröffentlicht: (2024)
von: Do, Thang, et al.
Veröffentlicht: (2024)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
von: Dereich, Steffen, et al.
Veröffentlicht: (2023)
von: Dereich, Steffen, et al.
Veröffentlicht: (2023)
Is ReLU Adversarially Robust?
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
von: Brännvall, Rickard, et al.
Veröffentlicht: (2023)
von: Brännvall, Rickard, et al.
Veröffentlicht: (2023)
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
von: Alexander, Michael
Veröffentlicht: (2026)
von: Alexander, Michael
Veröffentlicht: (2026)
Topological obstruction to the training of shallow ReLU neural networks
von: Nurisso, Marco, et al.
Veröffentlicht: (2024)
von: Nurisso, Marco, et al.
Veröffentlicht: (2024)
Deep Legendre Transform
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
Mathematical analysis of the gradients in deep learning
von: Dereich, Steffen, et al.
Veröffentlicht: (2025)
von: Dereich, Steffen, et al.
Veröffentlicht: (2025)
On the Curse of Memory in Recurrent Neural Networks: Approximation and Optimization Analysis
von: Li, Zhong, et al.
Veröffentlicht: (2020)
von: Li, Zhong, et al.
Veröffentlicht: (2020)
Stability properties of Minimal Gated Unit neural networks
von: De Carli, Stefano, et al.
Veröffentlicht: (2026)
von: De Carli, Stefano, et al.
Veröffentlicht: (2026)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
von: Do, Thang, et al.
Veröffentlicht: (2025)
von: Do, Thang, et al.
Veröffentlicht: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
Manifold limit for the training of shallow graph convolutional neural networks
von: Tengler, Johanna, et al.
Veröffentlicht: (2026)
von: Tengler, Johanna, et al.
Veröffentlicht: (2026)
Chaos-Free Networks are Stable Recurrent Neural Networks
von: De Carli, Stefano, et al.
Veröffentlicht: (2026)
von: De Carli, Stefano, et al.
Veröffentlicht: (2026)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
Data-induced multiscale losses and efficient multirate gradient descent schemes
von: He, Juncai, et al.
Veröffentlicht: (2024)
von: He, Juncai, et al.
Veröffentlicht: (2024)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for space-time solutions of semilinear partial differential equations
von: Ackermann, Julia, et al.
Veröffentlicht: (2024)
von: Ackermann, Julia, et al.
Veröffentlicht: (2024)
Near-optimal estimates for the $\ell^p$-Lipschitz constants of deep random ReLU neural networks
von: Dirksen, Sjoerd, et al.
Veröffentlicht: (2025)
von: Dirksen, Sjoerd, et al.
Veröffentlicht: (2025)
On the existence of optimal shallow feedforward networks with ReLU activation
von: Dereich, Steffen, et al.
Veröffentlicht: (2023)
von: Dereich, Steffen, et al.
Veröffentlicht: (2023)
PyEPO: A PyTorch-based End-to-End Predict-then-Optimize Library for Linear and Integer Programming
von: Tang, Bo, et al.
Veröffentlicht: (2022)
von: Tang, Bo, et al.
Veröffentlicht: (2022)
Deep Learning for Continuous-Time Stochastic Control with Jumps
von: Cheridito, Patrick, et al.
Veröffentlicht: (2025)
von: Cheridito, Patrick, et al.
Veröffentlicht: (2025)
Conservation Law Breaking at the Edge of Stability: A Spectral Theory of Non-Convex Neural Network Optimization
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
von: Medeiros, Daniel Nobrega
Veröffentlicht: (2026)
On the Principles of ReLU Networks with One Hidden Layer
von: Huang, Changcun
Veröffentlicht: (2024)
von: Huang, Changcun
Veröffentlicht: (2024)
Kourkoutas-Beta: A Sunspike-Driven Adam Optimizer with Desert Flair
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
von: Kassinos, Stavros C.
Veröffentlicht: (2025)
Physics Informed Differentiable Solvers for Learning Parametric Solution Manifolds in Heterogeneous Physical Systems
von: Panahi, Milad, et al.
Veröffentlicht: (2026)
von: Panahi, Milad, et al.
Veröffentlicht: (2026)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
von: Zhang, Tony, et al.
Veröffentlicht: (2025)
von: Zhang, Tony, et al.
Veröffentlicht: (2025)
Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
von: Sao, Piyush
Veröffentlicht: (2026)
von: Sao, Piyush
Veröffentlicht: (2026)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for Kolmogorov partial differential equations with Lipschitz nonlinearities in the $L^p$-sense
von: Ackermann, Julia, et al.
Veröffentlicht: (2023)
von: Ackermann, Julia, et al.
Veröffentlicht: (2023)
Neural network optimization strategies and the topography of the loss landscape
von: Yu, Jianneng, et al.
Veröffentlicht: (2026)
von: Yu, Jianneng, et al.
Veröffentlicht: (2026)
Topology and Geometry of the Learning Space of ReLU Networks: Connectivity and Singularities
von: Nurisso, Marco, et al.
Veröffentlicht: (2026)
von: Nurisso, Marco, et al.
Veröffentlicht: (2026)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
von: Vary, Simon, et al.
Veröffentlicht: (2024)
von: Vary, Simon, et al.
Veröffentlicht: (2024)
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
Step-Aware Residual-Guided Diffusion for EEG Spatial Super-Resolution
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
von: Liu, Hongjun, et al.
Veröffentlicht: (2025)
Rapid training of Hamiltonian graph networks using random features
von: Rahma, Atamert, et al.
Veröffentlicht: (2025)
von: Rahma, Atamert, et al.
Veröffentlicht: (2025)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
On Global and Local Convergence of Iterative Linear Quadratic Optimization Algorithms for Discrete Time Nonlinear Control
von: Roulet, Vincent, et al.
Veröffentlicht: (2022)
von: Roulet, Vincent, et al.
Veröffentlicht: (2022)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
von: Yamchote, Phaphontee, et al.
Veröffentlicht: (2025)
von: Yamchote, Phaphontee, et al.
Veröffentlicht: (2025)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
von: Kranz, Julian, et al.
Veröffentlicht: (2025)
von: Kranz, Julian, et al.
Veröffentlicht: (2025)
GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks
von: Salishev, Sergey, et al.
Veröffentlicht: (2025)
von: Salishev, Sergey, et al.
Veröffentlicht: (2025)
Error analysis for the deep Kolmogorov method
von: Cîmpean, Iulian, et al.
Veröffentlicht: (2025)
von: Cîmpean, Iulian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
von: Do, Thang, et al.
Veröffentlicht: (2024) -
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
von: Dereich, Steffen, et al.
Veröffentlicht: (2023) -
Is ReLU Adversarially Robust?
von: Sooksatra, Korn, et al.
Veröffentlicht: (2024) -
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
von: Brännvall, Rickard, et al.
Veröffentlicht: (2023) -
The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
von: Alexander, Michael
Veröffentlicht: (2026)