Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
Fuente:
arXiv
Saved in:
| Main Authors: | Arous, Gérard Ben, Erdogdu, Murat A., Vural, Nuri Mert, Wu, Denny |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pruning is Optimal for Learning Sparse Features in High-Dimensions
by: Vural, Nuri Mert, et al.
Published: (2024)
by: Vural, Nuri Mert, et al.
Published: (2024)
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025)
by: Ren, Yunwei, et al.
Published: (2025)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
by: Vural, Nuri Mert, et al.
Published: (2026)
by: Vural, Nuri Mert, et al.
Published: (2026)
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
by: Mousavi-Hosseini, Alireza, et al.
Published: (2024)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2024)
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)
When Do Transformers Outperform Feedforward and Recurrent Networks? A Statistical Perspective
by: Mousavi-Hosseini, Alireza, et al.
Published: (2025)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2025)
Stochastic gradient descent in high dimensions for multi-spiked tensor PCA
by: Arous, Gérard Ben, et al.
Published: (2024)
by: Arous, Gérard Ben, et al.
Published: (2024)
Local geometry of high-dimensional mixture models: Effective spectral theory and dynamical transitions
by: Arous, Gerard Ben, et al.
Published: (2025)
by: Arous, Gerard Ben, et al.
Published: (2025)
Langevin dynamics for high-dimensional optimization: the case of multi-spiked tensor PCA
by: Arous, Gérard Ben, et al.
Published: (2024)
by: Arous, Gérard Ben, et al.
Published: (2024)
Spectral alignment of stochastic gradient descent for high-dimensional classification tasks
by: Arous, Gerard Ben, et al.
Published: (2023)
by: Arous, Gerard Ben, et al.
Published: (2023)
Permutation recovery of spikes in noisy high-dimensional tensor estimation
by: Arous, Gérard Ben, et al.
Published: (2024)
by: Arous, Gérard Ben, et al.
Published: (2024)
Robust Feature Learning for Multi-Index Models in High Dimensions
by: Mousavi-Hosseini, Alireza, et al.
Published: (2024)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2024)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)
by: Lee, Jason D., et al.
Published: (2024)
Post-Training with Policy Gradients: Optimality and the Base Model Barrier
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
by: Mousavi-Hosseini, Alireza, et al.
Published: (2026)
Optimal Excess Risk Bounds for Empirical Risk Minimization on $p$-Norm Linear Regression
by: Hanchi, Ayoub El, et al.
Published: (2023)
by: Hanchi, Ayoub El, et al.
Published: (2023)
Learning smooth functions in high dimensions: from sparse polynomials to deep neural networks
by: Adcock, Ben, et al.
Published: (2024)
by: Adcock, Ben, et al.
Published: (2024)
On the Efficiency of ERM in Feature Learning
by: Hanchi, Ayoub El, et al.
Published: (2024)
by: Hanchi, Ayoub El, et al.
Published: (2024)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
by: Nitanda, Atsushi, et al.
Published: (2023)
by: Nitanda, Atsushi, et al.
Published: (2023)
Escape dynamics and implicit bias of one-pass SGD in overparameterized quadratic networks
by: Bocchi, Dario, et al.
Published: (2026)
by: Bocchi, Dario, et al.
Published: (2026)
A Geometric Analysis of PCA
by: Hanchi, Ayoub El, et al.
Published: (2025)
by: Hanchi, Ayoub El, et al.
Published: (2025)
Nonlinear spiked covariance matrices and signal propagation in deep neural networks
by: Wang, Zhichao, et al.
Published: (2024)
by: Wang, Zhichao, et al.
Published: (2024)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
by: Kovačević, Filip, et al.
Published: (2026)
by: Kovačević, Filip, et al.
Published: (2026)
Beyond Labeling Oracles: What does it mean to steal ML models?
by: Shafran, Avital, et al.
Published: (2023)
by: Shafran, Avital, et al.
Published: (2023)
Minimax Linear Regression under the Quantile Risk
by: Hanchi, Ayoub El, et al.
Published: (2024)
by: Hanchi, Ayoub El, et al.
Published: (2024)
Collective variables of neural networks: empirical time evolution and scaling laws
by: Tovey, Samuel, et al.
Published: (2024)
by: Tovey, Samuel, et al.
Published: (2024)
A Separation in Heavy-Tailed Sampling: Gaussian vs. Stable Oracles for Proximal Samplers
by: He, Ye, et al.
Published: (2024)
by: He, Ye, et al.
Published: (2024)
SGD method for entropy error function with smoothing l0 regularization for neural networks
by: Nguyen, Trong-Tuan, et al.
Published: (2024)
by: Nguyen, Trong-Tuan, et al.
Published: (2024)
Exploring the loss landscape of regularized neural networks via convex duality
by: Kim, Sungyoon, et al.
Published: (2024)
by: Kim, Sungyoon, et al.
Published: (2024)
Heavy-Tail Phenomenon in Decentralized SGD
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
by: Gurbuzbalaban, Mert, et al.
Published: (2022)
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
by: Liu, Toni J. B., et al.
Published: (2024)
by: Liu, Toni J. B., et al.
Published: (2024)
Broken neural scaling laws in materials science
by: Großmann, Max, et al.
Published: (2026)
by: Großmann, Max, et al.
Published: (2026)
Implicit regularization of deep residual networks towards neural ODEs
by: Marion, Pierre, et al.
Published: (2023)
by: Marion, Pierre, et al.
Published: (2023)
Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
Sampling from the Mean-Field Stationary Distribution
by: Kook, Yunbum, et al.
Published: (2024)
by: Kook, Yunbum, et al.
Published: (2024)
Analysis of Langevin Monte Carlo from Poincaré to Log-Sobolev
by: Chewi, Sinho, et al.
Published: (2021)
by: Chewi, Sinho, et al.
Published: (2021)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
Frequency-adaptive tensor neural networks for high-dimensional multi-scale problems
by: Huang, Jizu, et al.
Published: (2025)
by: Huang, Jizu, et al.
Published: (2025)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Negotiated Representations to Prevent Overfitting in Machine Learning Applications
by: Korhan, Nuri, et al.
Published: (2023)
by: Korhan, Nuri, et al.
Published: (2023)
Sparse deep neural networks for nonparametric estimation in high-dimensional sparse regression
by: Wu, Dongya, et al.
Published: (2024)
by: Wu, Dongya, et al.
Published: (2024)
Similar Items
-
Pruning is Optimal for Learning Sparse Features in High-Dimensions
by: Vural, Nuri Mert, et al.
Published: (2024) -
Emergence and scaling laws in SGD learning of shallow neural networks
by: Ren, Yunwei, et al.
Published: (2025) -
Learning to Recall with Transformers Beyond Orthogonal Embeddings
by: Vural, Nuri Mert, et al.
Published: (2026) -
Learning Multi-Index Models with Neural Networks via Mean-Field Langevin Dynamics
by: Mousavi-Hosseini, Alireza, et al.
Published: (2024) -
From Information to Generative Exponent: Learning Rate Induces Phase Transitions in SGD
by: Tsiolis, Konstantinos Christopher, et al.
Published: (2025)