SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network
Fuente:
arXiv
Guardado en:
| Autores principales: | Galanti, Tomer, Siegel, Zachary S., Gupte, Aparna, Poggio, Tomaso |
|---|---|
| Formato: | Preprint |
| Publicado: |
2022
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Formation of Representations in Neural Networks
por: Ziyin, Liu, et al.
Publicado: (2024)
por: Ziyin, Liu, et al.
Publicado: (2024)
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024)
On Generalization Bounds for Neural Networks with Low Rank Layers
por: Pinto, Andrea, et al.
Publicado: (2024)
por: Pinto, Andrea, et al.
Publicado: (2024)
Does Weight Decay Enhance Training Stability?
por: Saether, Marius, et al.
Publicado: (2026)
por: Saether, Marius, et al.
Publicado: (2026)
On efficiently computable functions, deep networks and sparse compositionality
por: Poggio, Tomaso
Publicado: (2025)
por: Poggio, Tomaso
Publicado: (2025)
Tool Building as a Path to "Superintelligence"
por: Koplow, David, et al.
Publicado: (2026)
por: Koplow, David, et al.
Publicado: (2026)
Learning Sparse Compositional Functions with Norm-Constrained Neural Networks
por: Huang, Shuo, et al.
Publicado: (2026)
por: Huang, Shuo, et al.
Publicado: (2026)
On the Power of Decision Trees in Auto-Regressive Language Modeling
por: Gan, Yulu, et al.
Publicado: (2024)
por: Gan, Yulu, et al.
Publicado: (2024)
Directional Neural Collapse Explains Few-Shot Transfer in Self-Supervised Learning
por: Luthra, Achleshwar, et al.
Publicado: (2026)
por: Luthra, Achleshwar, et al.
Publicado: (2026)
Ubiquity of Emergent Hebbian Dynamics in Regularized Learning
por: Koplow, David, et al.
Publicado: (2025)
por: Koplow, David, et al.
Publicado: (2025)
Hierarchical Reasoning Models: Perspectives and Misconceptions
por: Ge, Renee, et al.
Publicado: (2025)
por: Ge, Renee, et al.
Publicado: (2025)
On the Alignment Between Supervised and Self-Supervised Contrastive Learning
por: Luthra, Achleshwar, et al.
Publicado: (2025)
por: Luthra, Achleshwar, et al.
Publicado: (2025)
Scalable Principal-Agent Contract Design via Gradient-Based Optimization
por: Galanti, Tomer, et al.
Publicado: (2025)
por: Galanti, Tomer, et al.
Publicado: (2025)
Self-Supervised Contrastive Learning is Approximately Supervised Contrastive Learning
por: Luthra, Achleshwar, et al.
Publicado: (2025)
por: Luthra, Achleshwar, et al.
Publicado: (2025)
pAI/MSc: ML Theory Research with Humans on the Loop
por: Abdelmoneum, Mahmoud, et al.
Publicado: (2026)
por: Abdelmoneum, Mahmoud, et al.
Publicado: (2026)
From SGD to Spectra: A Theory of Neural Network Weight Dynamics
por: Olsen, Brian Richard, et al.
Publicado: (2025)
por: Olsen, Brian Richard, et al.
Publicado: (2025)
Does Feedback Alignment Work at Biological Timescales?
por: Bacvanski, Marc Gong, et al.
Publicado: (2025)
por: Bacvanski, Marc Gong, et al.
Publicado: (2025)
Topological Invariance and Breakdown in Learning
por: Yang, Yongyi, et al.
Publicado: (2025)
por: Yang, Yongyi, et al.
Publicado: (2025)
Sparse Linear Regression and Lattice Problems
por: Gupte, Aparna, et al.
Publicado: (2024)
por: Gupte, Aparna, et al.
Publicado: (2024)
Exact Mean Square Linear Stability Analysis for SGD
por: Mulayoff, Rotem, et al.
Publicado: (2023)
por: Mulayoff, Rotem, et al.
Publicado: (2023)
Agentic Systems as Boosting Weak Reasoning Models
por: Sunkaraneni, Varun, et al.
Publicado: (2026)
por: Sunkaraneni, Varun, et al.
Publicado: (2026)
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
por: Sadrtdinov, Ildus, et al.
Publicado: (2025)
por: Sadrtdinov, Ildus, et al.
Publicado: (2025)
The Fair Language Model Paradox
por: Pinto, Andrea, et al.
Publicado: (2024)
por: Pinto, Andrea, et al.
Publicado: (2024)
Iterative regularization in classification via hinge loss diagonal descent
por: Apidopoulos, Vassilis, et al.
Publicado: (2022)
por: Apidopoulos, Vassilis, et al.
Publicado: (2022)
LLM Priors for ERM over Programs
por: Singhal, Shivam, et al.
Publicado: (2025)
por: Singhal, Shivam, et al.
Publicado: (2025)
Unraveling Syntax: How Language Models Learn Context-Free Grammars
por: Schulz, Laura Ying, et al.
Publicado: (2025)
por: Schulz, Laura Ying, et al.
Publicado: (2025)
Too Sharp, Too Sure: When Calibration Follows Curvature
por: Morosini, Alessandro, et al.
Publicado: (2026)
por: Morosini, Alessandro, et al.
Publicado: (2026)
Heterosynaptic Circuits Are Universal Gradient Machines
por: Ziyin, Liu, et al.
Publicado: (2025)
por: Ziyin, Liu, et al.
Publicado: (2025)
Do Deep Networks Forget Initialization? A Forgetting-Time View of Practical Inductive Bias
por: Das, Mohua, et al.
Publicado: (2026)
por: Das, Mohua, et al.
Publicado: (2026)
Learning Multi-Index Models with Hyper-Kernel Ridge Regression
por: Huang, Shuo, et al.
Publicado: (2025)
por: Huang, Shuo, et al.
Publicado: (2025)
Information Filtering Networks: Theoretical Foundations, Generative Methodologies, and Real-World Applications
por: Aste, Tomaso
Publicado: (2025)
por: Aste, Tomaso
Publicado: (2025)
Parameter Symmetry Potentially Unifies Deep Learning Theory
por: Ziyin, Liu, et al.
Publicado: (2025)
por: Ziyin, Liu, et al.
Publicado: (2025)
Position: A Theory of Deep Learning Must Include Compositional Sparsity
por: Danhofer, David A., et al.
Publicado: (2025)
por: Danhofer, David A., et al.
Publicado: (2025)
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks
por: Kosson, Atli, et al.
Publicado: (2023)
por: Kosson, Atli, et al.
Publicado: (2023)
Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural Collapse
por: Jacot, Arthur, et al.
Publicado: (2024)
por: Jacot, Arthur, et al.
Publicado: (2024)
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks
por: Chen, Ke, et al.
Publicado: (2024)
por: Chen, Ke, et al.
Publicado: (2024)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
por: Attia, Amit, et al.
Publicado: (2025)
por: Attia, Amit, et al.
Publicado: (2025)
Same Error, Different Function: The Optimizer as an Implicit Prior in Financial Time Series
por: Cortesi, Federico Vittorio, et al.
Publicado: (2026)
por: Cortesi, Federico Vittorio, et al.
Publicado: (2026)
Momentum Further Constrains Sharpness at the Edge of Stochastic Stability
por: Andreyev, Arseniy, et al.
Publicado: (2026)
por: Andreyev, Arseniy, et al.
Publicado: (2026)
The Generalized Turing Test: A Foundation for Comparing Intelligence
por: Mitropolsky, Daniel, et al.
Publicado: (2026)
por: Mitropolsky, Daniel, et al.
Publicado: (2026)
Ejemplares similares
-
Formation of Representations in Neural Networks
por: Ziyin, Liu, et al.
Publicado: (2024) -
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD
por: Beneventano, Pierfrancesco, et al.
Publicado: (2024) -
On Generalization Bounds for Neural Networks with Low Rank Layers
por: Pinto, Andrea, et al.
Publicado: (2024) -
Does Weight Decay Enhance Training Stability?
por: Saether, Marius, et al.
Publicado: (2026) -
On efficiently computable functions, deep networks and sparse compositionality
por: Poggio, Tomaso
Publicado: (2025)