Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jacot, Arthur, Súkeník, Peter, Wang, Zihan, Mondelli, Marco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
von: Súkeník, Peter, et al.
Veröffentlicht: (2024)
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
von: Tucat, Matteo, et al.
Veröffentlicht: (2024)
von: Tucat, Matteo, et al.
Veröffentlicht: (2024)
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
von: Súkeník, Peter, et al.
Veröffentlicht: (2025)
von: Súkeník, Peter, et al.
Veröffentlicht: (2025)
Does Weight Decay Enhance Training Stability?
von: Saether, Marius, et al.
Veröffentlicht: (2026)
von: Saether, Marius, et al.
Veröffentlicht: (2026)
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
von: Laus, Hannah, et al.
Veröffentlicht: (2025)
Provably-Stable Neural Network-Based Control of Nonlinear Systems
von: Li, Anran, et al.
Veröffentlicht: (2025)
von: Li, Anran, et al.
Veröffentlicht: (2025)
Provable Acceleration of Nesterov's Accelerated Gradient for Rectangular Matrix Factorization and Linear Neural Networks
von: Xu, Zhenghao, et al.
Veröffentlicht: (2024)
von: Xu, Zhenghao, et al.
Veröffentlicht: (2024)
Cautious Weight Decay
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
von: Chen, Lizhang, et al.
Veröffentlicht: (2025)
Provable Accelerated Convergence of Nesterov's Momentum for Deep ReLU Neural Networks
von: Liao, Fangshuo, et al.
Veröffentlicht: (2023)
von: Liao, Fangshuo, et al.
Veröffentlicht: (2023)
Provable Acceleration of Nesterov's Accelerated Gradient Method over Heavy Ball Method in Training Over-Parameterized Neural Networks
von: Liu, Xin, et al.
Veröffentlicht: (2022)
von: Liu, Xin, et al.
Veröffentlicht: (2022)
Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable Data
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
von: Min, Hancheng, et al.
Veröffentlicht: (2025)
Weight-Parameterization in Continuous Time Deep Neural Networks for Surrogate Modeling
von: Rosso, Haley, et al.
Veröffentlicht: (2025)
von: Rosso, Haley, et al.
Veröffentlicht: (2025)
SMiLE: Provably Enforcing Global Relational Properties in Neural Networks
von: Francobaldi, Matteo, et al.
Veröffentlicht: (2025)
von: Francobaldi, Matteo, et al.
Veröffentlicht: (2025)
Relaxation-Informed Training of Neural Network Surrogate Models
von: Tsay, Calvin
Veröffentlicht: (2026)
von: Tsay, Calvin
Veröffentlicht: (2026)
Adaptive Momentum and Nonlinear Damping for Neural Network Training
von: Karoni, Aikaterini, et al.
Veröffentlicht: (2026)
von: Karoni, Aikaterini, et al.
Veröffentlicht: (2026)
Compression-aware Training of Neural Networks using Frank-Wolfe
von: Zimmer, Max, et al.
Veröffentlicht: (2022)
von: Zimmer, Max, et al.
Veröffentlicht: (2022)
Optimization Over Trained Neural Networks: Taking a Relaxing Walk
von: Tong, Jiatai, et al.
Veröffentlicht: (2024)
von: Tong, Jiatai, et al.
Veröffentlicht: (2024)
Convex Formulations for Training Two-Layer ReLU Neural Networks
von: Prakhya, Karthik, et al.
Veröffentlicht: (2024)
von: Prakhya, Karthik, et al.
Veröffentlicht: (2024)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
Towards Guided Descent: Optimization Algorithms for Training Neural Networks At Scale
von: Nagwekar, Ansh
Veröffentlicht: (2025)
von: Nagwekar, Ansh
Veröffentlicht: (2025)
Neural Network Training Techniques Regularize Optimization Trajectory: An Empirical Study
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
von: Chen, Cheng, et al.
Veröffentlicht: (2020)
Exploring the Potential of Bilevel Optimization for Calibrating Neural Networks
von: Sanguin, Gabriele, et al.
Veröffentlicht: (2025)
von: Sanguin, Gabriele, et al.
Veröffentlicht: (2025)
A Non-Monotone Preconditioned Trust-Region Method for Neural Network Training
von: Angino, Andrea, et al.
Veröffentlicht: (2026)
von: Angino, Andrea, et al.
Veröffentlicht: (2026)
Dual Natural Gradient Descent for Scalable Training of Physics-Informed Neural Networks
von: Jnini, Anas, et al.
Veröffentlicht: (2025)
von: Jnini, Anas, et al.
Veröffentlicht: (2025)
Neighbor-Sampling Based Momentum Stochastic Methods for Training Graph Neural Networks
von: Noel, Molly, et al.
Veröffentlicht: (2025)
von: Noel, Molly, et al.
Veröffentlicht: (2025)
Mean-Field Limits for Two-Layer Neural Networks Trained with Consensus-Based Optimization
von: De Deyn, William, et al.
Veröffentlicht: (2025)
von: De Deyn, William, et al.
Veröffentlicht: (2025)
Multi-Objective Linear Ensembles for Robust and Sparse Training of Few-Bit Neural Networks
von: Bernardelli, Ambrogio Maria, et al.
Veröffentlicht: (2022)
von: Bernardelli, Ambrogio Maria, et al.
Veröffentlicht: (2022)
Enhancing Stability of Physics-Informed Neural Network Training Through Saddle-Point Reformulation
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
von: Bylinkin, Dmitry, et al.
Veröffentlicht: (2025)
An Adaptive and Stability-Promoting Layerwise Training Approach for Sparse Deep Neural Network Architecture
von: Krishnanunni, C G, et al.
Veröffentlicht: (2022)
von: Krishnanunni, C G, et al.
Veröffentlicht: (2022)
Optimization over Trained (and Sparse) Neural Networks: A Surrogate within a Surrogate
von: Pham, Hung, et al.
Veröffentlicht: (2025)
von: Pham, Hung, et al.
Veröffentlicht: (2025)
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses
von: Tanguy, Eloi
Veröffentlicht: (2023)
von: Tanguy, Eloi
Veröffentlicht: (2023)
Scalable Solution of the Stochastic Multi-path Traveling Salesman Problem via Neural Networks
von: Chou, Xiaochen, et al.
Veröffentlicht: (2026)
von: Chou, Xiaochen, et al.
Veröffentlicht: (2026)
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
von: Yang, Yahong, et al.
Veröffentlicht: (2023)
von: Yang, Yahong, et al.
Veröffentlicht: (2023)
KKT-Informed Neural Network
von: Femine, Carmine Delle
Veröffentlicht: (2024)
von: Femine, Carmine Delle
Veröffentlicht: (2024)
Optimal Depth of Neural Networks
von: Qi, Qian
Veröffentlicht: (2025)
von: Qi, Qian
Veröffentlicht: (2025)
Muon is Provably Faster with Momentum Variance Reduction
von: Qian, Xun, et al.
Veröffentlicht: (2025)
von: Qian, Xun, et al.
Veröffentlicht: (2025)
Regularized Adaptive Momentum Dual Averaging with an Efficient Inexact Subproblem Solver for Training Structured Neural Network
von: Huang, Zih-Syuan, et al.
Veröffentlicht: (2024)
von: Huang, Zih-Syuan, et al.
Veröffentlicht: (2024)
Training Safe Neural Networks with Global SDP Bounds
von: Soletskyi, Roman, et al.
Veröffentlicht: (2024)
von: Soletskyi, Roman, et al.
Veröffentlicht: (2024)
Randomized Geometric Algebra Methods for Convex Neural Networks
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
von: Wang, Yifei, et al.
Veröffentlicht: (2024)
Rethinking the Capacity of Graph Neural Networks for Branching Strategy
von: Chen, Ziang, et al.
Veröffentlicht: (2024)
von: Chen, Ziang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Neural Collapse versus Low-rank Bias: Is Deep Neural Collapse Really Optimal?
von: Súkeník, Peter, et al.
Veröffentlicht: (2024) -
Regularized Gradient Clipping Provably Trains Wide and Deep Neural Networks
von: Tucat, Matteo, et al.
Veröffentlicht: (2024) -
Neural Collapse is Globally Optimal in Deep Regularized ResNets and Transformers
von: Súkeník, Peter, et al.
Veröffentlicht: (2025) -
Does Weight Decay Enhance Training Stability?
von: Saether, Marius, et al.
Veröffentlicht: (2026) -
Solving Inverse Problems with Deep Linear Neural Networks: Global Convergence Guarantees for Gradient Descent with Weight Decay
von: Laus, Hannah, et al.
Veröffentlicht: (2025)