A Theory of Non-Linear Feature Learning with One Gradient Step in Two-Layer Neural Networks
Fuente:
arXiv
Saved in:
| Main Authors: | Moniri, Behrad, Lee, Donghwan, Hassani, Hamed, Dobriban, Edgar |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026)
by: Moniri, Behrad, et al.
Published: (2026)
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
by: Moniri, Behrad, et al.
Published: (2025)
by: Moniri, Behrad, et al.
Published: (2025)
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024)
by: Moniri, Behrad, et al.
Published: (2024)
On The Concurrence of Layer-wise Preconditioning Methods and Provable Feature Learning
by: Zhang, Thomas T., et al.
Published: (2025)
by: Zhang, Thomas T., et al.
Published: (2025)
Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
by: Huang, Xinmeng, et al.
Published: (2023)
by: Huang, Xinmeng, et al.
Published: (2023)
Provable tradeoffs in adversarially robust classification
by: Dobriban, Edgar, et al.
Published: (2020)
by: Dobriban, Edgar, et al.
Published: (2020)
MultiRisk: Multiple Risk Control via Iterative Score Thresholding
by: Joshi, Sunay, et al.
Published: (2025)
by: Joshi, Sunay, et al.
Published: (2025)
Risk-Controlled Post-Processing of Decision Policies
by: Joshi, Sunay, et al.
Published: (2026)
by: Joshi, Sunay, et al.
Published: (2026)
Analysis of Off-Policy Multi-Step TD-Learning with Linear Function Approximation
by: Lee, Donghwan
Published: (2024)
by: Lee, Donghwan
Published: (2024)
Watermarking Language Models with Error Correcting Codes
by: Chao, Patrick, et al.
Published: (2024)
by: Chao, Patrick, et al.
Published: (2024)
Provable Multi-Task Representation Learning by Two-Layer ReLU Neural Networks
by: Collins, Liam, et al.
Published: (2023)
by: Collins, Liam, et al.
Published: (2023)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Conformal Inference under High-Dimensional Covariate Shifts via Likelihood-Ratio Regularization
by: Joshi, Sunay, et al.
Published: (2025)
by: Joshi, Sunay, et al.
Published: (2025)
A Switching System Theory of Q-Learning with Linear Function Approximation
by: Lee, Donghwan, et al.
Published: (2026)
by: Lee, Donghwan, et al.
Published: (2026)
Analysis of Off-Policy $n$-Step TD-Learning with Linear Function Approximation
by: Lim, Han-Dong, et al.
Published: (2025)
by: Lim, Han-Dong, et al.
Published: (2025)
How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
by: Dandi, Yatin, et al.
Published: (2023)
by: Dandi, Yatin, et al.
Published: (2023)
Conformal Information Pursuit for Interactively Guiding Large Language Models
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
by: Chan, Kwan Ho Ryan, et al.
Published: (2025)
Asymptotic Analysis of Two-Layer Neural Networks after One Gradient Step under Gaussian Mixtures Data with Structure
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Alternating Gradient Flows: A Theory of Feature Learning in Two-layer Neural Networks
by: Kunin, Daniel, et al.
Published: (2025)
by: Kunin, Daniel, et al.
Published: (2025)
Jailbreaking Black Box Large Language Models in Twenty Queries
by: Chao, Patrick, et al.
Published: (2023)
by: Chao, Patrick, et al.
Published: (2023)
Lyapunov-Certified Direct Switching Theory for Q-Learning
by: Lee, Donghwan
Published: (2026)
by: Lee, Donghwan
Published: (2026)
Statistical Methods in Generative AI
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Conformal Prediction with Learned Features
by: Kiyani, Shayan, et al.
Published: (2024)
by: Kiyani, Shayan, et al.
Published: (2024)
Solving a Research Problem in Mathematical Statistics with AI Assistance
by: Dobriban, Edgar
Published: (2025)
by: Dobriban, Edgar
Published: (2025)
Balancing Learning Rates Across Layers: Exact Two-Step Dynamics and Optimal Scaling in Linear Neural Networks
by: Pang, Tianyu, et al.
Published: (2026)
by: Pang, Tianyu, et al.
Published: (2026)
R-GTD: A Geometric Analysis of Gradient Temporal-Difference Learning in Singular Regimes
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
How Does Gradient Descent Learn Features -- A Local Analysis for Regularized Two-Layer Neural Networks
by: Zhou, Mo, et al.
Published: (2024)
by: Zhou, Mo, et al.
Published: (2024)
Uncertainty in Language Models: Assessment through Rank-Calibration
by: Huang, Xinmeng, et al.
Published: (2024)
by: Huang, Xinmeng, et al.
Published: (2024)
Optimal Decision-Making Based on Prediction Sets
by: Wang, Tao, et al.
Published: (2026)
by: Wang, Tao, et al.
Published: (2026)
Fair Classification by Direct Intervention on Operating Characteristics
by: Jiang, Kevin, et al.
Published: (2025)
by: Jiang, Kevin, et al.
Published: (2025)
Deep Q-Learning with Gradient Target Tracking
by: Park, Bum Geun, et al.
Published: (2025)
by: Park, Bum Geun, et al.
Published: (2025)
Soft Deterministic Policy Gradient with Gaussian Smoothing
by: Na, Hyunjun, et al.
Published: (2026)
by: Na, Hyunjun, et al.
Published: (2026)
Taming the Adversary: Stable Minimax Deep Deterministic Policy Gradient via Fractional Objectives
by: Lee, Taeho, et al.
Published: (2026)
by: Lee, Taeho, et al.
Published: (2026)
Minimax Statistical Estimation under Wasserstein Contamination
by: Chao, Patrick, et al.
Published: (2023)
by: Chao, Patrick, et al.
Published: (2023)
New Versions of Gradient Temporal Difference Learning
by: Lee, Donghwan, et al.
Published: (2021)
by: Lee, Donghwan, et al.
Published: (2021)
Bayes-Optimal Fair Classification with Linear Disparity Constraints via Pre-, In-, and Post-processing
by: Zeng, Xianli, et al.
Published: (2024)
by: Zeng, Xianli, et al.
Published: (2024)
SymmPI: Predictive Inference for Data with Group Symmetries
by: Dobriban, Edgar, et al.
Published: (2023)
by: Dobriban, Edgar, et al.
Published: (2023)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023)
by: Nichani, Eshaan, et al.
Published: (2023)
Similar Items
-
Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent
by: Moniri, Behrad, et al.
Published: (2026) -
Evaluating the Performance of Large Language Models via Debates
by: Moniri, Behrad, et al.
Published: (2024) -
Asymptotics of Linear Regression with Linearly Dependent Data
by: Moniri, Behrad, et al.
Published: (2024) -
On the Mechanisms of Weak-to-Strong Generalization: A Theoretical Perspective
by: Moniri, Behrad, et al.
Published: (2025) -
Signal-Plus-Noise Decomposition of Nonlinear Spiked Random Matrix Models
by: Moniri, Behrad, et al.
Published: (2024)