The Non-Linearity Perturbation Threshold: Width Scaling and Landscape Bifurcations in Deep Learning
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Alexander, Michael |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022)
PyEPO: A PyTorch-based End-to-End Predict-then-Optimize Library for Linear and Integer Programming
von: Tang, Bo, et al.
Veröffentlicht: (2022)
von: Tang, Bo, et al.
Veröffentlicht: (2022)
Multi-Objective Optimization with Desirability and Morris-Mitchell Criterion
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2025)
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2025)
i-DEQ: A stable inertial deep equilibrium model for image restoration
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
von: Clerc, Antonin, et al.
Veröffentlicht: (2026)
Interior-Point Vanishing Problem in Semidefinite Relaxations for Neural Network Verification
von: Ueda, Ryota, et al.
Veröffentlicht: (2025)
von: Ueda, Ryota, et al.
Veröffentlicht: (2025)
Ghosts of Softmax: Complex Singularities That Limit Safe Step Sizes in Cross-Entropy
von: Sao, Piyush
Veröffentlicht: (2026)
von: Sao, Piyush
Veröffentlicht: (2026)
Refining Graphical Neural Network Predictions Using Flow Matching for Optimal Power Flow with Constraint-Satisfaction Guarantee
von: Khanal, Kshitiz
Veröffentlicht: (2025)
von: Khanal, Kshitiz
Veröffentlicht: (2025)
Escaping Saddle Points via Curvature-Calibrated Perturbations: A Complete Analysis with Explicit Constants and Empirical Validation
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
von: Alpay, Faruk, et al.
Veröffentlicht: (2025)
Local properties of neural networks through the lens of layer-wise Hessians
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2025)
Inter-Layer Hessian Analysis of Neural Networks with DAG Architectures
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
Stochastic Estimation of the Layer-wise Hessian Trace for Monitoring Neural-network Training
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
von: Bolshim, Maxim, et al.
Veröffentlicht: (2026)
GDNSQ: Gradual Differentiable Noise Scale Quantization for Low-bit Neural Networks
von: Salishev, Sergey, et al.
Veröffentlicht: (2025)
von: Salishev, Sergey, et al.
Veröffentlicht: (2025)
From Non-Identifiability to Goal-Integrated Decision-Making in Parametric Inverse Optimization
von: Ahmadi, Farzin, et al.
Veröffentlicht: (2026)
von: Ahmadi, Farzin, et al.
Veröffentlicht: (2026)
Distributionally Robust Geometric Joint Chance-Constrained Optimization: Neurodynamic Approaches
von: Valli, Ange, et al.
Veröffentlicht: (2026)
von: Valli, Ange, et al.
Veröffentlicht: (2026)
Efficient Training of Physics-enhanced Neural ODEs via Direct Collocation and Nonlinear Programming
von: Langenkamp, Linus, et al.
Veröffentlicht: (2025)
von: Langenkamp, Linus, et al.
Veröffentlicht: (2025)
Sparse Training of Neural Networks based on Multilevel Mirror Descent
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
von: Lunk, Yannick, et al.
Veröffentlicht: (2026)
CAO: Curvature-Adaptive Optimization via Periodic Low-Rank Hessian Sketching
von: Du, Wenzhang
Veröffentlicht: (2025)
von: Du, Wenzhang
Veröffentlicht: (2025)
Enhancing Model Based Derivative Free Optimization using Direct Search
von: Li, Zijun, et al.
Veröffentlicht: (2026)
von: Li, Zijun, et al.
Veröffentlicht: (2026)
Black-Box Uniform Stability for Non-Euclidean Empirical Risk Minimization
von: Vary, Simon, et al.
Veröffentlicht: (2024)
von: Vary, Simon, et al.
Veröffentlicht: (2024)
Multi-Objective Optimization and Hyperparameter Tuning With Desirability Functions
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2025)
von: Bartz-Beielstein, Thomas
Veröffentlicht: (2025)
Adam Improves Muon: Adaptive Moment Estimation with Orthogonalized Momentum
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
von: Zhang, Minxin, et al.
Veröffentlicht: (2026)
Amazon Locker Capacity Management
von: Sethuraman, Samyukta, et al.
Veröffentlicht: (2023)
von: Sethuraman, Samyukta, et al.
Veröffentlicht: (2023)
Adaptive Conditional Forest Sampling for Spectral Risk Optimisation under Decision-Dependent Uncertainty
von: Kurbucz, Marcell T.
Veröffentlicht: (2026)
von: Kurbucz, Marcell T.
Veröffentlicht: (2026)
Deep Legendre Transform
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
von: Minabutdinov, Aleksey, et al.
Veröffentlicht: (2025)
On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning Regime
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
von: Jiang, Shuai, et al.
Veröffentlicht: (2026)
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
von: Vary, Simon, et al.
Veröffentlicht: (2026)
von: Vary, Simon, et al.
Veröffentlicht: (2026)
TOPSIS-like metaheuristic for LABS problem
von: Urbańczyk, Aleksandra, et al.
Veröffentlicht: (2025)
von: Urbańczyk, Aleksandra, et al.
Veröffentlicht: (2025)
prunAdag: an adaptive pruning-aware gradient method
von: Porcelli, Margherita, et al.
Veröffentlicht: (2025)
von: Porcelli, Margherita, et al.
Veröffentlicht: (2025)
Anisotropic Gaussian Smoothing for Gradient-based Optimization
von: Starnes, Andrew, et al.
Veröffentlicht: (2024)
von: Starnes, Andrew, et al.
Veröffentlicht: (2024)
Adaptive Methods for Multiobjective Unit Commitment
von: Tevruez, Ece, et al.
Veröffentlicht: (2025)
von: Tevruez, Ece, et al.
Veröffentlicht: (2025)
Last-iterate Convergence of ADMM on Multi-affine Quadratic Equality Constrained Problem
von: Chao, Yutong, et al.
Veröffentlicht: (2026)
von: Chao, Yutong, et al.
Veröffentlicht: (2026)
Manifold limit for the training of shallow graph convolutional neural networks
von: Tengler, Johanna, et al.
Veröffentlicht: (2026)
von: Tengler, Johanna, et al.
Veröffentlicht: (2026)
Adaptive Stochastic Gradient Descent Ascent Algorithm for Nonconvex Minimax Problems with Decision-Dependent Distributions
von: Gao, Yan, et al.
Veröffentlicht: (2025)
von: Gao, Yan, et al.
Veröffentlicht: (2025)
Where's Ben Nevis? A 2D optimisation benchmark with 957,174 local optima based on Great Britain terrain data
von: Wei, Yuhang, et al.
Veröffentlicht: (2024)
von: Wei, Yuhang, et al.
Veröffentlicht: (2024)
Universal Neural Optimal Transport
von: Geuter, Jonathan, et al.
Veröffentlicht: (2022)
von: Geuter, Jonathan, et al.
Veröffentlicht: (2022)
Setwise Coordinate Descent for Dual Asynchronous Decentralized Optimization
von: Costantini, Marina, et al.
Veröffentlicht: (2025)
von: Costantini, Marina, et al.
Veröffentlicht: (2025)
Implicit Bias and Invariance: How Hopfield Networks Efficiently Learn Graph Orbits
von: Murray, Michael, et al.
Veröffentlicht: (2025)
von: Murray, Michael, et al.
Veröffentlicht: (2025)
Machine Learning Algorithms for Improving Black Box Optimization Solvers
von: Kimiaei, Morteza, et al.
Veröffentlicht: (2025)
von: Kimiaei, Morteza, et al.
Veröffentlicht: (2025)
Quadratic Programming over Linearly Ordered Fields: Decidability and Attainment of Optimal Solutions
von: Plutenko, Dmytro O.
Veröffentlicht: (2026)
von: Plutenko, Dmytro O.
Veröffentlicht: (2026)
Deep Networks are Reproducing Kernel Chains
von: Heeringa, Tjeerd Jan, et al.
Veröffentlicht: (2025)
von: Heeringa, Tjeerd Jan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Gradient descent provably escapes saddle points in the training of shallow ReLU networks
von: Cheridito, Patrick, et al.
Veröffentlicht: (2022) -
PyEPO: A PyTorch-based End-to-End Predict-then-Optimize Library for Linear and Integer Programming
von: Tang, Bo, et al.
Veröffentlicht: (2022) -
Multi-Objective Optimization with Desirability and Morris-Mitchell Criterion
von: Bartz-Beielstein, Thomas, et al.
Veröffentlicht: (2025) -
i-DEQ: A stable inertial deep equilibrium model for image restoration
von: Clerc, Antonin, et al.
Veröffentlicht: (2026) -
Interior-Point Vanishing Problem in Semidefinite Relaxations for Neural Network Verification
von: Ueda, Ryota, et al.
Veröffentlicht: (2025)