Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Moniri, Behrad, Hassani, Hamed |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023)
by: Moniri, Behrad, et al.
Published: (2023)
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
by: Moniri, Behrad, et al.
Published: (2025)
by: Moniri, Behrad, et al.
Published: (2025)
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
by: Zhang, Thomas T., et al.
Published: (2025)
by: Zhang, Thomas T., et al.
Published: (2025)
Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
by: Collins, Liam, et al.
Published: (2023)
by: Collins, Liam, et al.
Published: (2023)
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
by: Zhou, Mo, et al.
Published: (2024)
by: Zhou, Mo, et al.
Published: (2024)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
Stochastic Gradient Descent for Two-layer Neural Networks
by: Cao, Dinghao, et al.
Published: (2024)
by: Cao, Dinghao, et al.
Published: (2024)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Convergence of Implicit Gradient Descent for Training Two-Layer Physics-Informed Neural Networks
by: Xu, Xianliang, et al.
Published: (2024)
by: Xu, Xianliang, et al.
Published: (2024)
The Benefits of Reusing Batches for Gradient Descent in Two-Layer Networks: Breaking the Curse of Information and Leap Exponents
by: Dandi, Yatin, et al.
Published: (2024)
by: Dandi, Yatin, et al.
Published: (2024)
Conformal Prediction with Learned Features
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
by: Cai, Yuhang, et al.
Published: (2024)
by: Cai, Yuhang, et al.
Published: (2024)
The Double Descent Behavior in Two Layer Neural Network for Binary Classification
by: Abeykoon, Chathurika S, et al.
Published: (2025)
by: Abeykoon, Chathurika S, et al.
Published: (2025)
Step by Step: Adaptive Gradient Descent for Training L-Lipschitz Neural Networks
by: Sung, Kyle, et al.
Published: (2025)
by: Sung, Kyle, et al.
Published: (2025)
On the Optimization and Generalization of Two-layer Transformers with Sign Gradient Descent
by: Li, Bingrui, et al.
Published: (2024)
by: Li, Bingrui, et al.
Published: (2024)
Homotopy Relaxation Training Algorithms for Infinite-Width Two-Layer ReLU Neural Networks
by: Yang, Yahong, et al.
Published: (2023)
by: Yang, Yahong, et al.
Published: (2023)
Comparing Spectral Bias and Robustness For Two-Layer Neural Networks: SGD vs Adaptive Random Fourier Features
by: Kammonen, Aku, et al.
Published: (2024)
by: Kammonen, Aku, et al.
Published: (2024)
Generalization Guarantees of Gradient Descent for Multi-Layer Neural Networks
by: Wang, Puyu, et al.
Published: (2023)
by: Wang, Puyu, et al.
Published: (2023)
One-Step Flow Policy Mirror Descent
by: Chen, Tianyi, et al.
Published: (2025)
by: Chen, Tianyi, et al.
Published: (2025)
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
by: Kunin, Daniel, et al.
Published: (2025)
by: Kunin, Daniel, et al.
Published: (2025)
Gradient Descent Converges Linearly to Flatter Minima than Gradient Flow in Shallow Linear Networks
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
by: Beneventano, Pierfrancesco, et al.
Published: (2025)
Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
by: Corlouer, Guillaume, et al.
Published: (2026)
by: Corlouer, Guillaume, et al.
Published: (2026)
Two-Step Q-Learning
by: Vijesh, Antony, et al.
Published: (2024)
by: Vijesh, Antony, et al.
Published: (2024)
Two-Timescale Gradient Descent Ascent Algorithms for Nonconvex Minimax Optimization
by: Lin, Tianyi, et al.
Published: (2024)
by: Lin, Tianyi, et al.
Published: (2024)
On the Sample Complexity of Two-Layer Networks: Lipschitz vs. Element-Wise Lipschitz Activation
by: Daniely, Amit, et al.
Published: (2022)
by: Daniely, Amit, et al.
Published: (2022)
Revisiting the Initial Steps in Adaptive Gradient Descent Optimization
by: Abuduweili, Abulikemu, et al.
Published: (2024)
by: Abuduweili, Abulikemu, et al.
Published: (2024)
Adaptive Step Sizes for Preconditioned Stochastic Gradient Descent
by: Köhne, Frederik, et al.
Published: (2023)
by: Köhne, Frederik, et al.
Published: (2023)
Simplicity Bias of Two-Layer Networks beyond Linearly Separable Data
by: Tsoy, Nikita, et al.
Published: (2024)
by: Tsoy, Nikita, et al.
Published: (2024)
Transfer Learning in Infinite Width Feature Learning Networks
by: Lauditi, Clarissa, et al.
Published: (2025)
by: Lauditi, Clarissa, et al.
Published: (2025)
Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
by: Cheng, Xiang, et al.
Published: (2023)
by: Cheng, Xiang, et al.
Published: (2023)
Absence of Closed-Form Descriptions for Gradient Flow in Two-Layer Narrow Networks
by: Park, Yeachan
Published: (2024)
by: Park, Yeachan
Published: (2024)
Gradient Descent with Large Step Sizes: Chaos and Fractal Convergence Region
by: Liang, Shuang, et al.
Published: (2025)
by: Liang, Shuang, et al.
Published: (2025)
Automated Feature Labeling with Token-Space Gradient Descent
by: Schulz, Julian, et al.
Published: (2025)
by: Schulz, Julian, et al.
Published: (2025)
Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks
by: Li, Binghui, et al.
Published: (2024)
by: Li, Binghui, et al.
Published: (2024)
Hybrid Coordinate Descent for Efficient Neural Network Learning Using Line Search and Gradient Descent
by: Hsiao, Yen-Che, et al.
Published: (2024)
by: Hsiao, Yen-Che, et al.
Published: (2024)
Federated Temporal Difference Learning with Linear Function Approximation under Environmental Heterogeneity
by: Wang, Han, et al.
Published: (2023)
by: Wang, Han, et al.
Published: (2023)
Similar Items
-
A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
by: Moniri, Behrad, et al.
Published: (2023) -
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024) -
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
by: Moniri, Behrad, et al.
Published: (2025) -
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024) -
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)