How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Mo, Ge, Rong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
von: Moniri, Behrad, et al.
Veröffentlicht: (2026)
von: Moniri, Behrad, et al.
Veröffentlicht: (2026)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
von: Wang, Puyu, et al.
Veröffentlicht: (2023)
von: Wang, Puyu, et al.
Veröffentlicht: (2023)
Stochastic Gradient Descent for Two-layer Neural Networks
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
von: Cao, Dinghao, et al.
Veröffentlicht: (2024)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
von: Moniri, Behrad, et al.
Veröffentlicht: (2023)
von: Moniri, Behrad, et al.
Veröffentlicht: (2023)
Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient Descent
von: Imai, Shota, et al.
Veröffentlicht: (2026)
von: Imai, Shota, et al.
Veröffentlicht: (2026)
The Double Descent Behavior in Two Layer Neural Network for Binary Classification
von: Abeykoon, Chathurika S, et al.
Veröffentlicht: (2025)
von: Abeykoon, Chathurika S, et al.
Veröffentlicht: (2025)
How Does Preconditioning Guide Feature Learning in Deep Neural Networks?
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
von: Yoshida, Kotaro, et al.
Veröffentlicht: (2025)
Locally Regularized Sparse Graph by Fast Proximal Gradient Descent
von: Sun, Dongfang, et al.
Veröffentlicht: (2024)
von: Sun, Dongfang, et al.
Veröffentlicht: (2024)
How Does the ReLU Activation Affect the Implicit Bias of Gradient Descent on High-dimensional Neural Network Regression?
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
von: Lai, Kuo-Wei, et al.
Veröffentlicht: (2026)
Data Diversity as Implicit Regularization: How Does Diversity Shape the Weight Space of Deep Neural Networks?
von: Ba, Yang, et al.
Veröffentlicht: (2024)
von: Ba, Yang, et al.
Veröffentlicht: (2024)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
von: Li, Binghui, et al.
Veröffentlicht: (2024)
von: Li, Binghui, et al.
Veröffentlicht: (2024)
On the Theory of Continual Learning with Gradient Descent for Neural Networks
von: Taheri, Hossein, et al.
Veröffentlicht: (2025)
von: Taheri, Hossein, et al.
Veröffentlicht: (2025)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
von: Hsiao, Yen-Che, et al.
Veröffentlicht: (2024)
von: Hsiao, Yen-Che, et al.
Veröffentlicht: (2024)
Convergence of Gradient Descent for Recurrent Neural Networks: A Nonasymptotic Analysis
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
von: Cayci, Semih, et al.
Veröffentlicht: (2024)
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
von: Dandi, Yatin, et al.
Veröffentlicht: (2024)
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
von: Kunin, Daniel, et al.
Veröffentlicht: (2025)
von: Kunin, Daniel, et al.
Veröffentlicht: (2025)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
Variational Stochastic Gradient Descent for Deep Neural Networks
von: Chen, Haotian, et al.
Veröffentlicht: (2024)
von: Chen, Haotian, et al.
Veröffentlicht: (2024)
Convergence Analysis of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
von: Xu, Xianliang, et al.
Veröffentlicht: (2024)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
von: Cai, Yuhang, et al.
Veröffentlicht: (2024)
von: Cai, Yuhang, et al.
Veröffentlicht: (2024)
Generalization Bounds of Stochastic Gradient Descent in Homogeneous Neural Networks
von: Ma, Wenquan, et al.
Veröffentlicht: (2026)
von: Ma, Wenquan, et al.
Veröffentlicht: (2026)
Distributed Gradient Descent for Functional Learning
von: Yu, Zhan, et al.
Veröffentlicht: (2023)
von: Yu, Zhan, et al.
Veröffentlicht: (2023)
Preconditioning for Accelerated Gradient Descent Optimization and Regularization
von: Ye, Qiang
Veröffentlicht: (2024)
von: Ye, Qiang
Veröffentlicht: (2024)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
von: Dandi, Yatin, et al.
Veröffentlicht: (2023)
Gradient Descent Robustly Learns the Intrinsic Dimension of Data in Training Convolutional Neural Networks
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2025)
How Transformers Learn Causal Structure with Gradient Descent
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
von: Nichani, Eshaan, et al.
Veröffentlicht: (2024)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
von: Li, Bingrui, et al.
Veröffentlicht: (2024)
von: Li, Bingrui, et al.
Veröffentlicht: (2024)
Do Neural Networks Need Gradient Descent to Generalize? A Theoretical Study
von: Alexander, Yotam, et al.
Veröffentlicht: (2025)
von: Alexander, Yotam, et al.
Veröffentlicht: (2025)
Automated Feature Labeling with Token-Space Gradient Descent
von: Schulz, Julian, et al.
Veröffentlicht: (2025)
von: Schulz, Julian, et al.
Veröffentlicht: (2025)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
von: Demir, Samet, et al.
Veröffentlicht: (2025)
von: Demir, Samet, et al.
Veröffentlicht: (2025)
Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural Networks
von: Peleg, Amit, et al.
Veröffentlicht: (2024)
von: Peleg, Amit, et al.
Veröffentlicht: (2024)
Recovery Guarantees of Unsupervised Neural Networks for Inverse Problems trained with Gradient Descent
von: Buskulic, Nathan, et al.
Veröffentlicht: (2024)
von: Buskulic, Nathan, et al.
Veröffentlicht: (2024)
Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
von: Sung, Kyle, et al.
Veröffentlicht: (2025)
von: Sung, Kyle, et al.
Veröffentlicht: (2025)
Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels
von: Yang, Jia-Qi, et al.
Veröffentlicht: (2025)
von: Yang, Jia-Qi, et al.
Veröffentlicht: (2025)
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
von: Sato, Naoki, et al.
Veröffentlicht: (2024)
Finite-Particle Rates for Regularized Stein Variational Gradient Descent
von: He, Ye, et al.
Veröffentlicht: (2026)
von: He, Ye, et al.
Veröffentlicht: (2026)
Understanding the Implicit Regularization of Gradient Descent in Over-parameterized Models
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
von: Ma, Jianhao, et al.
Veröffentlicht: (2025)
The Power of Random Features and the Limits of Distribution-Free Gradient Descent
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
von: Karchmer, Ari, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
von: Moniri, Behrad, et al.
Veröffentlicht: (2026) -
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
von: Wang, Puyu, et al.
Veröffentlicht: (2023) -
Stochastic Gradient Descent for Two-layer Neural Networks
von: Cao, Dinghao, et al.
Veröffentlicht: (2024) -
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
von: Xu, Xianliang, et al.
Veröffentlicht: (2024) -
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
von: Moniri, Behrad, et al.
Veröffentlicht: (2023)