Deconstructing the Goldilocks Zone of Neural Network Initialization
Fuente:
arXiv
Saved in:
| Main Authors: | Vysogorets, Artem, Dawid, Anna, Kempe, Julia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DRoP: Distributionally Robust Data Pruning
by: Vysogorets, Artem, et al.
Published: (2024)
by: Vysogorets, Artem, et al.
Published: (2024)
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
by: Kumar, Akshay, et al.
Published: (2024)
by: Kumar, Akshay, et al.
Published: (2024)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
by: Das, Mohua, et al.
Published: (2026)
by: Das, Mohua, et al.
Published: (2026)
KKT-Informed Neural Network
by: Femine, Carmine Delle
Published: (2024)
by: Femine, Carmine Delle
Published: (2024)
Optimal Depth of Neural Networks
by: Qi, Qian
Published: (2025)
by: Qi, Qian
Published: (2025)
Curse of Dimensionality in Neural Network Optimization
by: Na, Sanghoon, et al.
Published: (2025)
by: Na, Sanghoon, et al.
Published: (2025)
On the Topology of Neural Network Superlevel Sets
by: Gharesifard, Bahman
Published: (2026)
by: Gharesifard, Bahman
Published: (2026)
Learning Neural Networks by Neuron Pursuit
by: Kumar, Akshay, et al.
Published: (2025)
by: Kumar, Akshay, et al.
Published: (2025)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
by: Jacot, Arthur, et al.
Published: (2024)
by: Jacot, Arthur, et al.
Published: (2024)
A Recovery Guarantee for Sparse Neural Networks
by: Fridovich-Keil, Sara, et al.
Published: (2025)
by: Fridovich-Keil, Sara, et al.
Published: (2025)
Towards Quantifying the Hessian Structure of Neural Networks
by: Dong, Zhaorui, et al.
Published: (2025)
by: Dong, Zhaorui, et al.
Published: (2025)
A Unified Representation of Neural Networks Architectures
by: Prieur, Christophe, et al.
Published: (2025)
by: Prieur, Christophe, et al.
Published: (2025)
Convergence of Gradient Descent with Small Initialization for Unregularized Matrix Completion
by: Ma, Jianhao, et al.
Published: (2024)
by: Ma, Jianhao, et al.
Published: (2024)
Diagonalizing the Softmax: Hadamard Initialization for Tractable Cross-Entropy Dynamics
by: Garrod, Connall, et al.
Published: (2025)
by: Garrod, Connall, et al.
Published: (2025)
Forward Invariance in Neural Network Controlled Systems
by: Harapanahalli, Akash, et al.
Published: (2023)
by: Harapanahalli, Akash, et al.
Published: (2023)
Black-Box Approximation and Optimization with Hierarchical Tucker Decomposition
by: Ryzhakov, Gleb, et al.
Published: (2024)
by: Ryzhakov, Gleb, et al.
Published: (2024)
Gluon: Making Muon & Scion Great Again! (Bridging Theory and Practice of LMO-based Optimizers for LLMs)
by: Riabinin, Artem, et al.
Published: (2025)
by: Riabinin, Artem, et al.
Published: (2025)
Randomized Geometric Algebra Methods for Convex Neural Networks
by: Wang, Yifei, et al.
Published: (2024)
by: Wang, Yifei, et al.
Published: (2024)
Wasserstein Distributionally Robust Shallow Convex Neural Networks
by: Pallage, Julien, et al.
Published: (2024)
by: Pallage, Julien, et al.
Published: (2024)
SGD with Partial Hessian for Deep Neural Networks Optimization
by: Sun, Ying, et al.
Published: (2024)
by: Sun, Ying, et al.
Published: (2024)
Rethinking the Capacity of Graph Neural Networks for Branching Strategy
by: Chen, Ziang, et al.
Published: (2024)
by: Chen, Ziang, et al.
Published: (2024)
Regularized Gauss-Newton for Optimizing Overparameterized Neural Networks
by: Adeoye, Adeyemi D., et al.
Published: (2024)
by: Adeoye, Adeyemi D., et al.
Published: (2024)
Deep Operator Neural Network Model Predictive Control
by: de Jong, Thomas Oliver, et al.
Published: (2025)
by: de Jong, Thomas Oliver, et al.
Published: (2025)
Relaxation-Informed Training of Neural Network Surrogate Models
by: Tsay, Calvin
Published: (2026)
by: Tsay, Calvin
Published: (2026)
Exploring the Potential of Bilevel Optimization for Calibrating Neural Networks
by: Sanguin, Gabriele, et al.
Published: (2025)
by: Sanguin, Gabriele, et al.
Published: (2025)
Adaptive Momentum and Nonlinear Damping for Neural Network Training
by: Karoni, Aikaterini, et al.
Published: (2026)
by: Karoni, Aikaterini, et al.
Published: (2026)
On Integer Programming for the Binarized Neural Network Verification Problem
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
Chordal Sparsity for SDP-based Neural Network Verification
by: Xue, Anton, et al.
Published: (2022)
by: Xue, Anton, et al.
Published: (2022)
Distributed Control of Network Systems in the Space of Stabilizing Graph Neural Network Policies
by: Cao, John, et al.
Published: (2025)
by: Cao, John, et al.
Published: (2025)
Sample Complexity of Linear Quadratic Regulator Without Initial Stability
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2025)
by: Moghaddam, Amirreza Neshaei, et al.
Published: (2025)
Global Convergence of Four-Layer Matrix Factorization under Random Initialization
by: Luo, Minrui, et al.
Published: (2025)
by: Luo, Minrui, et al.
Published: (2025)
Improved Scalable Lipschitz Bounds for Deep Neural Networks
by: Syed, Usman, et al.
Published: (2025)
by: Syed, Usman, et al.
Published: (2025)
Physics-Informed Neural Networks with Hard Linear Equality Constraints
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Loss Landscape Characterization of Neural Networks without Over-Parametrization
by: Islamov, Rustem, et al.
Published: (2024)
by: Islamov, Rustem, et al.
Published: (2024)
Compression-aware Training of Neural Networks using Frank-Wolfe
by: Zimmer, Max, et al.
Published: (2022)
by: Zimmer, Max, et al.
Published: (2022)
SensLI: Sensitivity-Based Layer Insertion for Neural Networks
by: Kreis, Leonie, et al.
Published: (2023)
by: Kreis, Leonie, et al.
Published: (2023)
Optimality-Informed Neural Networks for Solving Parametric Optimization Problems
by: Hoffmann, Matthias K., et al.
Published: (2025)
by: Hoffmann, Matthias K., et al.
Published: (2025)
FSNet: Feasibility-Seeking Neural Network for Constrained Optimization with Guarantees
by: Nguyen, Hoang T., et al.
Published: (2025)
by: Nguyen, Hoang T., et al.
Published: (2025)
Robust Angular Synchronization via Directed Graph Neural Networks
by: He, Yixuan, et al.
Published: (2023)
by: He, Yixuan, et al.
Published: (2023)
Similar Items
-
DRoP: Distributionally Robust Data Pruning
by: Vysogorets, Artem, et al.
Published: (2024) -
Early Directional Convergence in Deep Homogeneous Neural Networks for Small Initializations
by: Kumar, Akshay, et al.
Published: (2024) -
Directional Convergence Near Small Initializations and Saddles in Two-Homogeneous Neural Networks
by: Kumar, Akshay, et al.
Published: (2024) -
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
by: Das, Mohua, et al.
Published: (2026) -
KKT-Informed Neural Network
by: Femine, Carmine Delle
Published: (2024)