Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
Fuente:
arXiv
Saved in:
| Main Authors: | Alexander, Yotam, Slutzky, Yonatan, Ran-Milo, Yuval, Cohen, Nadav |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels
by: Slutzky, Yonatan, et al.
Published: (2024)
by: Slutzky, Yonatan, et al.
Published: (2024)
Why Does Agentic Safety Fail to Generalize Across Tasks?
by: Slutzky, Yonatan, et al.
Published: (2026)
by: Slutzky, Yonatan, et al.
Published: (2026)
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
Mamba Knockout for Unraveling Factual Information Flow
by: Endy, Nir, et al.
Published: (2025)
by: Endy, Nir, et al.
Published: (2025)
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026)
by: Ran-Milo, Yuval
Published: (2026)
What Makes Data Suitable for a Locally Connected Neural Network? A Necessary and Sufficient Condition Based on Quantum Entanglement
by: Alexander, Yotam, et al.
Published: (2023)
by: Alexander, Yotam, et al.
Published: (2023)
Provable Benefits of Complex Parameterizations for Structured State Space Models
by: Ran-Milo, Yuval, et al.
Published: (2024)
by: Ran-Milo, Yuval, et al.
Published: (2024)
Lecture Notes on Linear Neural Networks: A Tale of Optimization and Generalization in Deep Learning
by: Cohen, Nadav, et al.
Published: (2024)
by: Cohen, Nadav, et al.
Published: (2024)
Implicit Bias of Policy Gradient in Linear Quadratic Control: Extrapolation to Unseen Initial States
by: Razin, Noam, et al.
Published: (2024)
by: Razin, Noam, et al.
Published: (2024)
A Mechanistic Account of Attention Sinks in GPT-2: One Circuit, Broader Implications for Mitigation
by: Ran-Milo, Yuval, et al.
Published: (2026)
by: Ran-Milo, Yuval, et al.
Published: (2026)
FSW-GNN: A Bi-Lipschitz WL-Equivalent Graph Neural Network
by: Sverdlov, Yonatan, et al.
Published: (2024)
by: Sverdlov, Yonatan, et al.
Published: (2024)
In-context Learning and Gradient Descent Revisited
by: Deutch, Gilad, et al.
Published: (2023)
by: Deutch, Gilad, et al.
Published: (2023)
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
by: Ma, Wenquan, et al.
Published: (2026)
by: Ma, Wenquan, et al.
Published: (2026)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
by: Wang, Puyu, et al.
Published: (2023)
by: Wang, Puyu, et al.
Published: (2023)
On the Expressive Power of Sparse Geometric MPNNs
by: Sverdlov, Yonatan, et al.
Published: (2024)
by: Sverdlov, Yonatan, et al.
Published: (2024)
Stochastic Gradient Descent for Two-layer Neural Networks
by: Cao, Dinghao, et al.
Published: (2024)
by: Cao, Dinghao, et al.
Published: (2024)
Variational Stochastic Gradient Descent for Deep Neural Networks
by: Chen, Haotian, et al.
Published: (2024)
by: Chen, Haotian, et al.
Published: (2024)
Information-Theoretic Generalization Bounds for Stochastic Gradient Descent with Predictable Virtual Noise
by: Partohaghighi, Mohammad
Published: (2026)
by: Partohaghighi, Mohammad
Published: (2026)
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent
by: Wang, Mingze, et al.
Published: (2023)
by: Wang, Mingze, et al.
Published: (2023)
Monotone and Separable Set Functions: Characterizations and Neural Models
by: Sarangi, Soutrik, et al.
Published: (2025)
by: Sarangi, Soutrik, et al.
Published: (2025)
When and How to Canonize: A Generalization Perspective
by: Sverdlov, Yonatan, et al.
Published: (2026)
by: Sverdlov, Yonatan, et al.
Published: (2026)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
by: Hsiao, Yen-Che, et al.
Published: (2024)
by: Hsiao, Yen-Che, et al.
Published: (2024)
On the Theory of Continual Learning with Gradient Descent for Neural Networks
by: Taheri, Hossein, et al.
Published: (2025)
by: Taheri, Hossein, et al.
Published: (2025)
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
by: Corlouer, Guillaume, et al.
Published: (2026)
by: Corlouer, Guillaume, et al.
Published: (2026)
How Do the Architecture and Optimizer Affect Representation Learning? On the Training Dynamics of Representations in Deep Neural Networks
by: Sharon, Yuval, et al.
Published: (2024)
by: Sharon, Yuval, et al.
Published: (2024)
Gradient Descent Finds Over-Parameterized Neural Networks with Sharp Generalization for Nonparametric Regression
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
by: Cayci, Semih, et al.
Published: (2024)
by: Cayci, Semih, et al.
Published: (2024)
On the Hölder Stability of Multiset and Graph Neural Networks
by: Davidson, Yair, et al.
Published: (2024)
by: Davidson, Yair, et al.
Published: (2024)
Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
by: Sung, Kyle, et al.
Published: (2025)
by: Sung, Kyle, et al.
Published: (2025)
Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural Networks
by: Peleg, Amit, et al.
Published: (2024)
by: Peleg, Amit, et al.
Published: (2024)
Recovery Guarantees of Unsupervised Neural Networks for Inverse Problems trained with Gradient Descent
by: Buskulic, Nathan, et al.
Published: (2024)
by: Buskulic, Nathan, et al.
Published: (2024)
Revisiting Multi-Permutation Equivariance through the Lens of Irreducible Representations
by: Sverdlov, Yonatan, et al.
Published: (2024)
by: Sverdlov, Yonatan, et al.
Published: (2024)
On the Generalization of Stochastic Gradient Descent with Momentum
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
by: Ramezani-Kebrya, Ali, et al.
Published: (2018)
Non-Euclidean Gradient Descent Operates at the Edge of Stability
by: Islamov, Rustem, et al.
Published: (2026)
by: Islamov, Rustem, et al.
Published: (2026)
Condensed Stein Variational Gradient Descent for Uncertainty Quantification of Neural Networks
by: Padmanabha, Govinda Anantha, et al.
Published: (2024)
by: Padmanabha, Govinda Anantha, et al.
Published: (2024)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
by: Zhang, Chenyang, et al.
Published: (2025)
by: Zhang, Chenyang, et al.
Published: (2025)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Convergence Analysis of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
Approximation and Gradient Descent Training with Neural Networks
by: Welper, G.
Published: (2024)
by: Welper, G.
Published: (2024)
Growing Neural Networks: Dynamic Evolution through Gradient Descent
by: Radhakrishnan, Anil, et al.
Published: (2025)
by: Radhakrishnan, Anil, et al.
Published: (2025)
Similar Items
-
The Implicit Bias of Structured State Space Models Can Be Poisoned With Clean Labels
by: Slutzky, Yonatan, et al.
Published: (2024) -
Why Does Agentic Safety Fail to Generalize Across Tasks?
by: Slutzky, Yonatan, et al.
Published: (2026) -
Outcome-Based RL Provably Leads Transformers to Reason, but Only With the Right Data
by: Ran-Milo, Yuval, et al.
Published: (2026) -
Mamba Knockout for Unraveling Factual Information Flow
by: Endy, Nir, et al.
Published: (2025) -
Attention Sinks Are Provably Necessary in Softmax Transformers: Evidence from Trigger-Conditional Tasks
by: Ran-Milo, Yuval
Published: (2026)