Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
Fuente:
arXiv
Saved in:
| Main Authors: | Bantzis, Ioannis, Simon, James B., Jacot, Arthur |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
by: Wu, Frank Zhengqing, et al.
Published: (2024)
by: Wu, Frank Zhengqing, et al.
Published: (2024)
Dimer-Enhanced Optimization: A First-Order Approach to Escaping Saddle Points in Neural Network Training
by: Hu, Yue, et al.
Published: (2025)
by: Hu, Yue, et al.
Published: (2025)
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025)
by: Zhang, Yedi, et al.
Published: (2025)
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
by: Corlouer, Guillaume, et al.
Published: (2026)
by: Corlouer, Guillaume, et al.
Published: (2026)
A Theory of Saddle Escape in Deep Nonlinear Networks
by: Rawal, Divit, et al.
Published: (2026)
by: Rawal, Divit, et al.
Published: (2026)
N-ReLU: Zero-Mean Stochastic Extension of ReLU
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
by: Manik, Md Motaleb Hossen, et al.
Published: (2025)
The Geometry of ReLU Networks through the ReLU Transition Graph
by: Dhayalkar, Sahil Rajesh
Published: (2025)
by: Dhayalkar, Sahil Rajesh
Published: (2025)
The Resurrection of the ReLU
by: Horuz, Coşku Can, et al.
Published: (2025)
by: Horuz, Coşku Can, et al.
Published: (2025)
Dimension-Free Saddle-Point Escape in Muon
by: Long, Yanlin, et al.
Published: (2026)
by: Long, Yanlin, et al.
Published: (2026)
Only Strict Saddles in the Energy Landscape of Predictive Coding Networks?
by: Innocenti, Francesco, et al.
Published: (2024)
by: Innocenti, Francesco, et al.
Published: (2024)
Pathwise Explanation of ReLU Neural Networks
by: Lim, Seongwoo, et al.
Published: (2025)
by: Lim, Seongwoo, et al.
Published: (2025)
Beyond ReLU: Chebyshev-DQN for Enhanced Deep Q-Networks
by: Yazdannik, Saman, et al.
Published: (2025)
by: Yazdannik, Saman, et al.
Published: (2025)
Bilevel Optimization over Saddle Points of Zero-Sum Markov Games
by: Zheng, Zihao, et al.
Published: (2026)
by: Zheng, Zihao, et al.
Published: (2026)
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
by: Harzli, Ouns El, et al.
Published: (2026)
by: Harzli, Ouns El, et al.
Published: (2026)
Efficiently Escaping Saddle Points for Policy Optimization
by: Khorasani, Sadegh, et al.
Published: (2023)
by: Khorasani, Sadegh, et al.
Published: (2023)
Neural Network-based High-index Saddle Dynamics Method for Searching Saddle Points and Solution Landscape
by: Liu, Yuankai, et al.
Published: (2024)
by: Liu, Yuankai, et al.
Published: (2024)
Unveiling the Training Dynamics of ReLU Networks through a Linear Lens
by: Ye, Longqing
Published: (2025)
by: Ye, Longqing
Published: (2025)
Three Quantization Regimes for ReLU Networks
by: Ou, Weigutian, et al.
Published: (2024)
by: Ou, Weigutian, et al.
Published: (2024)
$λ$-GELU: Learning Gating Hardness for Controlled ReLU-ization in Deep Networks
by: Pérez-Corral, Cristian, et al.
Published: (2026)
by: Pérez-Corral, Cristian, et al.
Published: (2026)
Activation-Descent Regularization for Input Optimization of ReLU Networks
by: Yu, Hongzhan, et al.
Published: (2024)
by: Yu, Hongzhan, et al.
Published: (2024)
Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff
by: Jacot, Arthur
Published: (2023)
by: Jacot, Arthur
Published: (2023)
Relating Piecewise Linear Kolmogorov Arnold Networks to ReLU Networks
by: Schoots, Nandi, et al.
Published: (2025)
by: Schoots, Nandi, et al.
Published: (2025)
ReLU Networks for Exact Generation of Similar Graphs
by: Ghafoor, Mamoona, et al.
Published: (2026)
by: Ghafoor, Mamoona, et al.
Published: (2026)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
by: Park, Junhyung, et al.
Published: (2024)
by: Park, Junhyung, et al.
Published: (2024)
Deep ReLU Networks Have Surprisingly Simple Polytopes
by: Fan, Feng-Lei, et al.
Published: (2023)
by: Fan, Feng-Lei, et al.
Published: (2023)
Convergence of Shallow ReLU Networks on Weakly Interacting Data
by: Dana, Léo, et al.
Published: (2025)
by: Dana, Léo, et al.
Published: (2025)
Expressive Power of ReLU and Step Networks under Floating-Point Operations
by: Park, Yeachan, et al.
Published: (2024)
by: Park, Yeachan, et al.
Published: (2024)
Is ReLU Adversarially Robust?
by: Sooksatra, Korn, et al.
Published: (2024)
by: Sooksatra, Korn, et al.
Published: (2024)
Geometry-induced Regularization in Deep ReLU Neural Networks
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
by: Bona-Pellissier, Joachim, et al.
Published: (2024)
RePO: Understanding Preference Learning Through ReLU-Based Optimization
by: Wu, Junkang, et al.
Published: (2025)
by: Wu, Junkang, et al.
Published: (2025)
Hessian-guided Perturbed Wasserstein Gradient Flows for Escaping Saddle Points
by: Yamamoto, Naoya, et al.
Published: (2025)
by: Yamamoto, Naoya, et al.
Published: (2025)
Uncertainty Quantification with Bayesian Higher Order ReLU KANs
by: Giroux, James, et al.
Published: (2024)
by: Giroux, James, et al.
Published: (2024)
Precise Verification of Transformers through ReLU-Catalyzed Abstraction Refinement
by: Liu, Hengjie, et al.
Published: (2026)
by: Liu, Hengjie, et al.
Published: (2026)
Uncovering Layer-Dependent Activation Sparsity Patterns in ReLU Transformers
by: Wild, Cody, et al.
Published: (2024)
by: Wild, Cody, et al.
Published: (2024)
Detecting Invariant Manifolds in ReLU-Based RNNs
by: Eisenmann, Lukas, et al.
Published: (2025)
by: Eisenmann, Lukas, et al.
Published: (2025)
Mixed Dynamics In Linear Networks: Unifying the Lazy and Active Regimes
by: Tu, Zhenfeng, et al.
Published: (2024)
by: Tu, Zhenfeng, et al.
Published: (2024)
Topological Signatures of ReLU Neural Network Activation Patterns
by: Bosca, Vicente, et al.
Published: (2025)
by: Bosca, Vicente, et al.
Published: (2025)
A Lower Bound for the Number of Linear Regions of Ternary ReLU Regression Neural Networks
by: Nakahara, Yuta, et al.
Published: (2025)
by: Nakahara, Yuta, et al.
Published: (2025)
Does Flatness imply Generalization for Logistic Loss in Univariate Two-Layer ReLU Network?
by: Qiao, Dan, et al.
Published: (2025)
by: Qiao, Dan, et al.
Published: (2025)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
Similar Items
-
Loss Landscape of Shallow ReLU-like Neural Networks: Stationary Points, Saddle Escape, and Network Embedding
by: Wu, Frank Zhengqing, et al.
Published: (2024) -
Dimer-Enhanced Optimization: A First-Order Approach to Escaping Saddle Points in Neural Network Training
by: Hu, Yue, et al.
Published: (2025) -
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
by: Zhang, Yedi, et al.
Published: (2025) -
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
by: Corlouer, Guillaume, et al.
Published: (2026) -
A Theory of Saddle Escape in Deep Nonlinear Networks
by: Rawal, Divit, et al.
Published: (2026)