Neural network learns low-dimensional polynomials with SGD near the information-theoretic limit
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Jason D., Oko, Kazusato, Suzuki, Taiji, Wu, Denny |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pretrained transformer efficiently learns low-dimensional target functions in-context
por: Oko, Kazusato, et al.
Publicado: (2024)
por: Oko, Kazusato, et al.
Publicado: (2024)
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
por: Oko, Kazusato, et al.
Publicado: (2024)
por: Oko, Kazusato, et al.
Publicado: (2024)
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Flow matching achieves almost minimax optimal convergence
por: Fukumizu, Kenji, et al.
Publicado: (2024)
por: Fukumizu, Kenji, et al.
Publicado: (2024)
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
por: Kim, Juno, et al.
Publicado: (2023)
por: Kim, Juno, et al.
Publicado: (2023)
Emergence and scaling laws in SGD learning of shallow neural networks
por: Ren, Yunwei, et al.
Publicado: (2025)
por: Ren, Yunwei, et al.
Publicado: (2025)
Metastable Dynamics of Chain-of-Thought Reasoning: Provable Benefits of Search, RL and Distillation
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
A Statistical Theory of Contrastive Pre-training and Multimodal Generative AI
por: Oko, Kazusato, et al.
Publicado: (2025)
por: Oko, Kazusato, et al.
Publicado: (2025)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Test time training enhances in-context learning of nonlinear functions
por: Kuwataka, Kento, et al.
Publicado: (2025)
por: Kuwataka, Kento, et al.
Publicado: (2025)
AutoLL: Automatic Linear Layout of Graphs based on Deep Neural Network
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
Mean-field Analysis on Two-layer Neural Networks from a Kernel Perspective
por: Takakura, Shokichi, et al.
Publicado: (2024)
por: Takakura, Shokichi, et al.
Publicado: (2024)
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
por: Arous, Gérard Ben, et al.
Publicado: (2025)
por: Arous, Gérard Ben, et al.
Publicado: (2025)
Dimensionality-induced information loss of outliers in deep neural networks
por: Uematsu, Kazuki, et al.
Publicado: (2024)
por: Uematsu, Kazuki, et al.
Publicado: (2024)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
por: Nitanda, Atsushi, et al.
Publicado: (2023)
por: Nitanda, Atsushi, et al.
Publicado: (2023)
On the Learning Dynamics of Two-layer Linear Networks with Label Noise SGD
por: Zhang, Tongcheng, et al.
Publicado: (2026)
por: Zhang, Tongcheng, et al.
Publicado: (2026)
How Does Label Noise Gradient Descent Improve Generalization in the Low SNR Regime?
por: Huang, Wei, et al.
Publicado: (2025)
por: Huang, Wei, et al.
Publicado: (2025)
From Shortcut to Induction Head: How Data Diversity Shapes Algorithm Selection in Transformers
por: Kawata, Ryotaro, et al.
Publicado: (2025)
por: Kawata, Ryotaro, et al.
Publicado: (2025)
Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
State Space Models are Provably Comparable to Transformers in Dynamic Token Selection
por: Nishikawa, Naoki, et al.
Publicado: (2024)
por: Nishikawa, Naoki, et al.
Publicado: (2024)
Transformers Provably Solve Parity Efficiently with Chain of Thought
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Deep Two-Way Matrix Reordering for Relational Data Analysis
por: Watanabe, Chihiro, et al.
Publicado: (2021)
por: Watanabe, Chihiro, et al.
Publicado: (2021)
The Mechanism of Weak-to-Strong Generalization: Feature Elicitation from Latent Knowledge
por: Awano, Ryoya, et al.
Publicado: (2026)
por: Awano, Ryoya, et al.
Publicado: (2026)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
por: Kawata, Ryotaro, et al.
Publicado: (2026)
por: Kawata, Ryotaro, et al.
Publicado: (2026)
In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-Learning
por: Wakayama, Tomoya, et al.
Publicado: (2025)
por: Wakayama, Tomoya, et al.
Publicado: (2025)
Approximation and Estimation Ability of Transformers for Sequence-to-Sequence Functions with Infinite Dimensional Input
por: Takakura, Shokichi, et al.
Publicado: (2023)
por: Takakura, Shokichi, et al.
Publicado: (2023)
High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
por: Jagannath, Aukosh, et al.
Publicado: (2025)
por: Jagannath, Aukosh, et al.
Publicado: (2025)
Direct Density Ratio Optimization: A Statistically Consistent Approach to Aligning Large Language Models
por: Higuchi, Rei, et al.
Publicado: (2025)
por: Higuchi, Rei, et al.
Publicado: (2025)
Towards a Unified Analysis of Neural Networks in Nonparametric Instrumental Variable Regression: Optimization and Generalization
por: Chen, Zonghao, et al.
Publicado: (2025)
por: Chen, Zonghao, et al.
Publicado: (2025)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
por: Kovačević, Filip, et al.
Publicado: (2026)
por: Kovačević, Filip, et al.
Publicado: (2026)
Optimality and Adaptivity of Deep Neural Features for Instrumental Variable Regression
por: Kim, Juno, et al.
Publicado: (2025)
por: Kim, Juno, et al.
Publicado: (2025)
Transformers are Minimax Optimal Nonparametric In-Context Learners
por: Kim, Juno, et al.
Publicado: (2024)
por: Kim, Juno, et al.
Publicado: (2024)
Degrees of Freedom for Linear Attention: Distilling Softmax Attention with Optimal Feature Efficiency
por: Nishikawa, Naoki, et al.
Publicado: (2025)
por: Nishikawa, Naoki, et al.
Publicado: (2025)
Mamba Can Learn Low-Dimensional Targets In-Context via Test-Time Feature Learning
por: Oh, Junsoo, et al.
Publicado: (2025)
por: Oh, Junsoo, et al.
Publicado: (2025)
From Saddle Points Toward Global Minima: A Newton-Type Method on Wasserstein Space
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
por: Lascu, Razvan-Andrei, et al.
Publicado: (2026)
Quantifying the Optimization and Generalization Advantages of Graph Neural Networks Over Multilayer Perceptrons
por: Huang, Wei, et al.
Publicado: (2023)
por: Huang, Wei, et al.
Publicado: (2023)
Intrinsic Wasserstein Rates for Score-Based Generative Models on Smooth Manifolds
por: Fu, Guoji, et al.
Publicado: (2026)
por: Fu, Guoji, et al.
Publicado: (2026)
Shallow diffusion networks provably learn hidden low-dimensional structure
por: Boffi, Nicholas M., et al.
Publicado: (2024)
por: Boffi, Nicholas M., et al.
Publicado: (2024)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
por: Abbe, Emmanuel, et al.
Publicado: (2022)
por: Abbe, Emmanuel, et al.
Publicado: (2022)
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory
por: Kim, Juno, et al.
Publicado: (2026)
por: Kim, Juno, et al.
Publicado: (2026)
Ejemplares similares
-
Pretrained transformer efficiently learns low-dimensional target functions in-context
por: Oko, Kazusato, et al.
Publicado: (2024) -
Learning sum of diverse features: computational hardness and efficient gradient-based training for ridge combinations
por: Oko, Kazusato, et al.
Publicado: (2024) -
Direct Distributional Optimization for Provable Alignment of Diffusion Models
por: Kawata, Ryotaro, et al.
Publicado: (2025) -
Flow matching achieves almost minimax optimal convergence
por: Fukumizu, Kenji, et al.
Publicado: (2024) -
Symmetric Mean-field Langevin Dynamics for Distributional Minimax Problems
por: Kim, Juno, et al.
Publicado: (2023)