On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Jentzen, Arnulf, Kassing, Sebastian |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
by: Kranz, Julian, et al.
Published: (2025)
by: Kranz, Julian, et al.
Published: (2025)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for space-time solutions of semilinear partial differential equations
by: Ackermann, Julia, et al.
Published: (2024)
by: Ackermann, Julia, et al.
Published: (2024)
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for Kolmogorov partial differential equations with Lipschitz nonlinearities in the $L^p$-sense
by: Ackermann, Julia, et al.
Published: (2023)
by: Ackermann, Julia, et al.
Published: (2023)
On the algorithmic construction of deep ReLU networks
by: Huybrechs, Daan
Published: (2025)
by: Huybrechs, Daan
Published: (2025)
Equidistribution-based training of Free Knot Splines and ReLU Neural Networks
by: Appella, Simone, et al.
Published: (2024)
by: Appella, Simone, et al.
Published: (2024)
ReLU neural network approximation to piecewise constant functions
by: Cai, Zhiqiang, et al.
Published: (2024)
by: Cai, Zhiqiang, et al.
Published: (2024)
Discretization Error of Fourier Neural Operators
by: Lanthaler, Samuel, et al.
Published: (2024)
by: Lanthaler, Samuel, et al.
Published: (2024)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
by: Cheridito, Patrick, et al.
Published: (2022)
by: Cheridito, Patrick, et al.
Published: (2022)
Near-optimal learning of Banach-valued, high-dimensional functions via deep neural networks
by: Adcock, Ben, et al.
Published: (2022)
by: Adcock, Ben, et al.
Published: (2022)
Mathematical analysis of the gradients in deep learning
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Fourier Residual Networks Achieve Spectral Accuracy for Discontinuous Functions
by: Davis, Owen, et al.
Published: (2026)
by: Davis, Owen, et al.
Published: (2026)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
Exact Loop Controllers for ReLU Realization of Homogeneous Curve Refinements
by: Bolorkhuu, Boldsaikhan, et al.
Published: (2026)
by: Bolorkhuu, Boldsaikhan, et al.
Published: (2026)
Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer
by: Li, Mingda, et al.
Published: (2024)
by: Li, Mingda, et al.
Published: (2024)
InhibiDistilbert: Knowledge Distillation for a ReLU and Addition-based Transformer
by: Zhang, Tony, et al.
Published: (2025)
by: Zhang, Tony, et al.
Published: (2025)
Topological obstruction to the training of shallow ReLU neural networks
by: Nurisso, Marco, et al.
Published: (2024)
by: Nurisso, Marco, et al.
Published: (2024)
On the Principles of ReLU Networks with One Hidden Layer
by: Huang, Changcun
Published: (2024)
by: Huang, Changcun
Published: (2024)
Adaptive Randomized Neural Networks with Locally Activation Function: Theory and Algorithm for Solving PDEs
by: Bi, Ran, et al.
Published: (2026)
by: Bi, Ran, et al.
Published: (2026)
Approximation bounds for norm constrained deep neural networks
by: Maiale, Francesco Paolo, et al.
Published: (2025)
by: Maiale, Francesco Paolo, et al.
Published: (2025)
Trustworthy AI in numerics: On verification algorithms for neural network-based PDE solvers
by: Haugen, Emil, et al.
Published: (2025)
by: Haugen, Emil, et al.
Published: (2025)
Reduced Order Modeling of Partial Differential Equations on Parameter-Dependent Domains Using Deep Neural Networks
by: Bukač, Martina, et al.
Published: (2024)
by: Bukač, Martina, et al.
Published: (2024)
An $r$-adaptive finite element method using neural networks for parametric self-adjoint elliptic problem
by: Aballay, Danilo, et al.
Published: (2025)
by: Aballay, Danilo, et al.
Published: (2025)
Generation of maximal snake polyominoes using a deep neural network
by: Gauthier, Benjamin, et al.
Published: (2026)
by: Gauthier, Benjamin, et al.
Published: (2026)
Parameter-Efficient Transformer Embeddings
by: Ndubuaku, Henry, et al.
Published: (2025)
by: Ndubuaku, Henry, et al.
Published: (2025)
DYNAMAX: Dynamic computing for Transformers and Mamba based architectures
by: Nogales, Miguel, et al.
Published: (2025)
by: Nogales, Miguel, et al.
Published: (2025)
Minimum Width of Leaky-ReLU Neural Networks for Uniform Universal Approximation
by: Li, Li'ang, et al.
Published: (2023)
by: Li, Li'ang, et al.
Published: (2023)
In almost all shallow analytic neural network optimization landscapes, efficient minimizers have strongly convex neighborhoods
by: Benning, Felix, et al.
Published: (2025)
by: Benning, Felix, et al.
Published: (2025)
Proof-Carrying Verification for ReLU Networks via Rational Certificates
by: Gokavarapu, Chandrasekhar
Published: (2025)
by: Gokavarapu, Chandrasekhar
Published: (2025)
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Understanding the geometry of deep learning with decision boundary volume
by: Burfitt, Matthew, et al.
Published: (2026)
by: Burfitt, Matthew, et al.
Published: (2026)
Exponential convergence rates for momentum stochastic gradient descent in the overparametrized setting
by: Gess, Benjamin, et al.
Published: (2023)
by: Gess, Benjamin, et al.
Published: (2023)
An adaptive adjoint-oriented neural network for solving parametric optimal control problems with singularities
by: Yuan, Zikang, et al.
Published: (2025)
by: Yuan, Zikang, et al.
Published: (2025)
Allure of Craquelure: A Variational-Generative Approach to Crack Detection in Paintings
by: Paul, Laura, et al.
Published: (2026)
by: Paul, Laura, et al.
Published: (2026)
Approximation Error and Complexity Bounds for ReLU Networks on Low-Regular Function Spaces
by: Davis, Owen, et al.
Published: (2024)
by: Davis, Owen, et al.
Published: (2024)
Exact ReLU realization of tensor-product refinement iterates
by: Gantumur, Tsogtgerel
Published: (2026)
by: Gantumur, Tsogtgerel
Published: (2026)
A Survey on Collaborating Small and Large Language Models for Performance, Cost-effectiveness, Cloud-edge Privacy, and Trustworthiness
by: Wang, Fali, et al.
Published: (2025)
by: Wang, Fali, et al.
Published: (2025)
Sharp higher order convergence rates for the Adam optimizer
by: Dereich, Steffen, et al.
Published: (2025)
by: Dereich, Steffen, et al.
Published: (2025)
Exact Sequence Interpolation with Transformers
by: Alcalde, Albert, et al.
Published: (2025)
by: Alcalde, Albert, et al.
Published: (2025)
Similar Items
-
On the existence of optimal shallow feedforward networks with ReLU activation
by: Dereich, Steffen, et al.
Published: (2023) -
SAD Neural Networks: Divergent Gradient Flows and Asymptotic Optimality via o-minimal Structures
by: Kranz, Julian, et al.
Published: (2025) -
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for space-time solutions of semilinear partial differential equations
by: Ackermann, Julia, et al.
Published: (2024) -
Asymptotic stability properties and a priori bounds for Adam and other gradient descent optimization methods
by: Dereich, Steffen, et al.
Published: (2025) -
Deep neural networks with ReLU, leaky ReLU, and softplus activation provably overcome the curse of dimensionality for Kolmogorov partial differential equations with Lipschitz nonlinearities in the $L^p$-sense
by: Ackermann, Julia, et al.
Published: (2023)