Emergence and scaling laws in SGD learning of shallow neural networks
Fuente:
arXiv
Saved in:
| Main Authors: | Ren, Yunwei, Nichani, Eshaan, Wu, Denny, Lee, Jason D. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026)
by: Kim, Juno, et al.
Published: (2026)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023)
by: Nichani, Eshaan, et al.
Published: (2023)
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
by: Arous, Gérard Ben, et al.
Published: (2025)
by: Arous, Gérard Ben, et al.
Published: (2025)
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Understanding Factual Recall in Transformers via Associative Memories
by: Nichani, Eshaan, et al.
Published: (2024)
by: Nichani, Eshaan, et al.
Published: (2024)
Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
by: Lee, Jason D., et al.
Published: (2024)
by: Lee, Jason D., et al.
Published: (2024)
On the Statistical Query Complexity of Learning Semiautomata: a Random Walk Approach
by: Giapitzakis, George, et al.
Published: (2025)
by: Giapitzakis, George, et al.
Published: (2025)
Fine-Tuning Dynamics of In-Context Factual Recall in Transformers
by: Huang, Ruomin, et al.
Published: (2026)
by: Huang, Ruomin, et al.
Published: (2026)
Learning Compositional Functions with Transformers from Easy-to-Hard Data
by: Wang, Zixuan, et al.
Published: (2025)
by: Wang, Zixuan, et al.
Published: (2025)
Learning Hierarchical Polynomials of Multiple Nonlinear Features with Three-Layer Networks
by: Fu, Hengyu, et al.
Published: (2024)
by: Fu, Hengyu, et al.
Published: (2024)
Sharp Capacity Thresholds in Linear Associative Memory: From Winner-Take-All to Listwise Retrieval
by: Barnfield, Nicholas, et al.
Published: (2026)
by: Barnfield, Nicholas, et al.
Published: (2026)
Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Fine-Tuning Language Models with Just Forward Passes
by: Malladi, Sadhika, et al.
Published: (2023)
by: Malladi, Sadhika, et al.
Published: (2023)
Orthogonal greedy algorithm for linear operator learning with shallow neural network
by: Lin, Ye, et al.
Published: (2025)
by: Lin, Ye, et al.
Published: (2025)
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers
by: Ren, Yunwei, et al.
Published: (2024)
by: Ren, Yunwei, et al.
Published: (2024)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
by: Nitanda, Atsushi, et al.
Published: (2023)
by: Nitanda, Atsushi, et al.
Published: (2023)
Provable Learning of Random Hierarchy Models and Hierarchical Shallow-to-Deep Chaining
by: Ren, Yunwei, et al.
Published: (2026)
by: Ren, Yunwei, et al.
Published: (2026)
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
by: D'Amico, Francesco, et al.
Published: (2025)
by: D'Amico, Francesco, et al.
Published: (2025)
Quantum advantage for learning shallow neural networks with natural data distributions
by: Lewis, Laura, et al.
Published: (2025)
by: Lewis, Laura, et al.
Published: (2025)
Nonlinear spiked covariance matrices and signal propagation in deep neural networks
by: Wang, Zhichao, et al.
Published: (2024)
by: Wang, Zhichao, et al.
Published: (2024)
Representations learnt by SGD and Adaptive learning rules: Conditions that vary sparsity and selectivity in neural networks
by: Park, Jin Hyun
Published: (2022)
by: Park, Jin Hyun
Published: (2022)
Collective variables of neural networks: empirical time evolution and scaling laws
by: Tovey, Samuel, et al.
Published: (2024)
by: Tovey, Samuel, et al.
Published: (2024)
Asymptotic convexity of wide and shallow neural networks
by: Borkar, Vivek, et al.
Published: (2025)
by: Borkar, Vivek, et al.
Published: (2025)
Are neural scaling laws leading quantum chemistry astray?
by: Lee, Siwoo, et al.
Published: (2025)
by: Lee, Siwoo, et al.
Published: (2025)
SGD method for entropy error function with smoothing l0 regularization for neural networks
by: Nguyen, Trong-Tuan, et al.
Published: (2024)
by: Nguyen, Trong-Tuan, et al.
Published: (2024)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
by: Kovačević, Filip, et al.
Published: (2026)
by: Kovačević, Filip, et al.
Published: (2026)
Sharp feature-learning transitions and Bayes-optimal neural scaling laws in extensive-width networks
by: Nguyen, Minh-Toan, et al.
Published: (2026)
by: Nguyen, Minh-Toan, et al.
Published: (2026)
LLMs learn governing principles of dynamical systems, revealing an in-context neural scaling law
by: Liu, Toni J. B., et al.
Published: (2024)
by: Liu, Toni J. B., et al.
Published: (2024)
Broken neural scaling laws in materials science
by: Großmann, Max, et al.
Published: (2026)
by: Großmann, Max, et al.
Published: (2026)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
by: Abbe, Emmanuel, et al.
Published: (2022)
by: Abbe, Emmanuel, et al.
Published: (2022)
On shallow feedforward neural networks with inputs from a topological space
by: Ismailov, Vugar
Published: (2025)
by: Ismailov, Vugar
Published: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
by: Kim, Juno, et al.
Published: (2025)
by: Kim, Juno, et al.
Published: (2025)
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD
by: Arnaboldi, Luca, et al.
Published: (2023)
by: Arnaboldi, Luca, et al.
Published: (2023)
Optimal generalisation and learning transition in extensive-width shallow neural networks near interpolation
by: Barbier, Jean, et al.
Published: (2025)
by: Barbier, Jean, et al.
Published: (2025)
Optimal rates of approximation by shallow ReLU$^k$ neural networks and applications to nonparametric regression
by: Yang, Yunfei, et al.
Published: (2023)
by: Yang, Yunfei, et al.
Published: (2023)
Nonparametric regression using over-parameterized shallow ReLU neural networks
by: Yang, Yunfei, et al.
Published: (2023)
by: Yang, Yunfei, et al.
Published: (2023)
Bayesian continual learning and forgetting in neural networks
by: Bonnet, Djohan, et al.
Published: (2025)
by: Bonnet, Djohan, et al.
Published: (2025)
Optimal scaling laws in learning hierarchical multi-index models
by: Defilippis, Leonardo, et al.
Published: (2026)
by: Defilippis, Leonardo, et al.
Published: (2026)
A shallow physics-informed neural network for solving partial differential equations on surfaces
by: Hu, Wei-Fan, et al.
Published: (2022)
by: Hu, Wei-Fan, et al.
Published: (2022)
Similar Items
-
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
by: Kim, Juno, et al.
Published: (2026) -
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025) -
Provable Guarantees for Nonlinear Feature Learning in Three-Layer Neural Networks
by: Nichani, Eshaan, et al.
Published: (2023) -
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
by: Arous, Gérard Ben, et al.
Published: (2025) -
How Transformers Learn Causal Structure with Gradient Descent
by: Nichani, Eshaan, et al.
Published: (2024)