Mathematical analysis of the gradients in deep learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dereich, Steffen, Do, Thang, Jentzen, Arnulf, Weber, Frederic |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error analysis for the deep Kolmogorov method
by: Cîmpean, Iulian, et al.
Published: (2025)
by: Cîmpean, Iulian, et al.
Published: (2025)
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025)
by: Do, Thang, et al.
Published: (2025)
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023)
by: Dereich, Steffen, et al.
Published: (2023)
The Manifold Scattering Transform for High-Dimensional Point Cloud Data
by: Chew, Joyce, et al.
Published: (2022)
by: Chew, Joyce, et al.
Published: (2022)
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)
by: Dereich, Steffen, et al.
Published: (2026)
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation
by: Do, Thang, et al.
Published: (2024)
by: Do, Thang, et al.
Published: (2024)
Neural Green's Operators for Parametric Partial Differential Equations
by: Melchers, Hugo, et al.
Published: (2024)
by: Melchers, Hugo, et al.
Published: (2024)
Mathematical artificial data for operator learning
by: Wu, Heng, et al.
Published: (2025)
by: Wu, Heng, et al.
Published: (2025)
Deep learning four decades of human migration
by: Gaskin, Thomas, et al.
Published: (2025)
by: Gaskin, Thomas, et al.
Published: (2025)
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
by: Cheridito, Patrick, et al.
Published: (2022)
by: Cheridito, Patrick, et al.
Published: (2022)
Your contrastive learning problem is secretly a distribution alignment problem
by: Chen, Zihao, et al.
Published: (2025)
by: Chen, Zihao, et al.
Published: (2025)
Exploring specialization and sensitivity of convolutional neural networks in the context of simultaneous image augmentations
by: Kharyuk, Pavel, et al.
Published: (2025)
by: Kharyuk, Pavel, et al.
Published: (2025)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Sprecher Networks: A Parameter-Efficient Kolmogorov-Arnold Architecture
by: Hägg, Christian, et al.
Published: (2025)
by: Hägg, Christian, et al.
Published: (2025)
Data-induced multiscale losses and efficient multirate gradient descent schemes
by: He, Juncai, et al.
Published: (2024)
by: He, Juncai, et al.
Published: (2024)
Application of Sensitivity Analysis Methods for Studying Neural Network Models
by: Miao, Jiaxuan, et al.
Published: (2025)
by: Miao, Jiaxuan, et al.
Published: (2025)
Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory
by: Jentzen, Arnulf, et al.
Published: (2023)
by: Jentzen, Arnulf, et al.
Published: (2023)
Deep Hedging Under Non-Convexity: Limitations and a Case for AlphaZero
by: Maggiolo, Matteo, et al.
Published: (2025)
by: Maggiolo, Matteo, et al.
Published: (2025)
Stage-wise Dynamics of Classifier-Free Guidance in Diffusion Models
by: Jin, Cheng, et al.
Published: (2025)
by: Jin, Cheng, et al.
Published: (2025)
FeNeC: Enhancing Continual Learning via Feature Clustering with Neighbor- or Logit-Based Classification
by: Książek, Kamil, et al.
Published: (2025)
by: Książek, Kamil, et al.
Published: (2025)
X-Factor: Quality Is a Dataset-Intrinsic Property
by: Couch, Josiah, et al.
Published: (2025)
by: Couch, Josiah, et al.
Published: (2025)
A Constraint-Preserving Neural Network Approach for Solving Mean-Field Games Equilibrium
by: Liu, Jinwei, et al.
Published: (2025)
by: Liu, Jinwei, et al.
Published: (2025)
The Domain Mixed Unit: A New Neural Arithmetic Layer
by: Curry, Paul
Published: (2025)
by: Curry, Paul
Published: (2025)
Drift-Resilient TabPFN: In-Context Learning Temporal Distribution Shifts on Tabular Data
by: Helli, Kai, et al.
Published: (2024)
by: Helli, Kai, et al.
Published: (2024)
IntSeqBERT: Learning Arithmetic Structure in OEIS via Modulo-Spectrum Embeddings
by: Nakasho, Kazuhisa
Published: (2026)
by: Nakasho, Kazuhisa
Published: (2026)
Developing Explainable Machine Learning Model using Augmented Concept Activation Vector
by: Hassanpour, Reza, et al.
Published: (2024)
by: Hassanpour, Reza, et al.
Published: (2024)
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
by: Dereich, Steffen, et al.
Published: (2024)
by: Dereich, Steffen, et al.
Published: (2024)
Just In Time Transformers
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks
by: Salishev, Sergey, et al.
Published: (2025)
by: Salishev, Sergey, et al.
Published: (2025)
Deep learning algorithms for solving high dimensional nonlinear backward stochastic differential equations
by: Kapllani, Lorenc, et al.
Published: (2020)
by: Kapllani, Lorenc, et al.
Published: (2020)
Physics-based deep kernel learning for parameter estimation in high dimensional PDEs
by: Yan, Weihao, et al.
Published: (2025)
by: Yan, Weihao, et al.
Published: (2025)
From Features to Graphs: Exploring Graph Structures and Pairwise Interactions via GNNs
by: Yamchote, Phaphontee, et al.
Published: (2025)
by: Yamchote, Phaphontee, et al.
Published: (2025)
Correction and Corruption: A Two-Rate View of Error Flow in LLM Protocols
by: Reitich, Fernando
Published: (2026)
by: Reitich, Fernando
Published: (2026)
Distinguished In Uniform: Self Attention Vs. Virtual Nodes
by: Rosenbluth, Eran, et al.
Published: (2024)
by: Rosenbluth, Eran, et al.
Published: (2024)
Enhancing Predictive Accuracy in Tennis: Integrating Fuzzy Logic and CV-GRNN for Dynamic Match Outcome and Player Momentum Analysis
by: Li, Kechen, et al.
Published: (2025)
by: Li, Kechen, et al.
Published: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Territory Paint Wars: Diagnosing and Mitigating Failure Modes in Competitive Multi-Agent PPO
by: Singh, Diyansha
Published: (2026)
by: Singh, Diyansha
Published: (2026)
Time Series Predictions in Unmonitored Sites: A Survey of Machine Learning Techniques in Water Resources
by: Willard, Jared D., et al.
Published: (2023)
by: Willard, Jared D., et al.
Published: (2023)
Similar Items
-
Error analysis for the deep Kolmogorov method
by: Cîmpean, Iulian, et al.
Published: (2025) -
Non-convergence to the optimal risk for Adam and stochastic gradient descent optimization in the training of deep neural networks
by: Do, Thang, et al.
Published: (2025) -
On the existence of minimizers in shallow residual ReLU neural network optimization landscapes
by: Dereich, Steffen, et al.
Published: (2023) -
The Manifold Scattering Transform for High-Dimensional Point Cloud Data
by: Chew, Joyce, et al.
Published: (2022) -
Uniform a priori bounds and error analysis for the Adam stochastic gradient descent optimization method
by: Dereich, Steffen, et al.
Published: (2026)