Limitations of SGD for Multi-Index Models Beyond Statistical Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Barzilai, Daniel, Shamir, Ohad |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Models Don't Collapse: On the Consistency of Iterative MLE
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Generalization in Kernel Regression Under Realistic Assumptions
by: Barzilai, Daniel, et al.
Published: (2023)
by: Barzilai, Daniel, et al.
Published: (2023)
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Simple Relative Deviation Bounds for Covariance and Gram Matrices
by: Barzilai, Daniel, et al.
Published: (2024)
by: Barzilai, Daniel, et al.
Published: (2024)
Are Convex Optimization Curves Convex?
by: Barzilai, Guy, et al.
Published: (2025)
by: Barzilai, Guy, et al.
Published: (2025)
Hardness of Learning Fixed Parities with Neural Networks
by: Shoshani, Itamar, et al.
Published: (2025)
by: Shoshani, Itamar, et al.
Published: (2025)
Gradient Descent's Last Iterate is Often (slightly) Suboptimal
by: Kornowski, Guy, et al.
Published: (2026)
by: Kornowski, Guy, et al.
Published: (2026)
An Algorithm with Optimal Dimension-Dependence for Zero-Order Nonsmooth Nonconvex Stochastic Optimization
by: Kornowski, Guy, et al.
Published: (2023)
by: Kornowski, Guy, et al.
Published: (2023)
On the Complexity of Finding Small Subgradients in Nonsmooth Optimization
by: Kornowski, Guy, et al.
Published: (2022)
by: Kornowski, Guy, et al.
Published: (2022)
Open Problem: Anytime Convergence Rate of Gradient Descent
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
Logarithmic Width Suffices for Robust Memorization
by: Egosi, Amitsour, et al.
Published: (2025)
by: Egosi, Amitsour, et al.
Published: (2025)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
The Oracle Complexity of Simplex-based Matrix Games
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
REED-VAE: RE-Encode Decode Training for Iterative Image Editing with Diffusion Models
by: Almog, Gal, et al.
Published: (2025)
by: Almog, Gal, et al.
Published: (2025)
On the Hardness of Meaningful Local Guarantees in Nonsmooth Nonconvex Optimization
by: Kornowski, Guy, et al.
Published: (2024)
by: Kornowski, Guy, et al.
Published: (2024)
From Tempered to Benign Overfitting in ReLU Neural Networks
by: Kornowski, Guy, et al.
Published: (2023)
by: Kornowski, Guy, et al.
Published: (2023)
Querying Kernel Methods Suffices for Reconstructing their Training Data
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Depth Separation in Norm-Bounded Infinite-Width Neural Networks
by: Parkinson, Suzanna, et al.
Published: (2024)
by: Parkinson, Suzanna, et al.
Published: (2024)
SLowcal-SGD: Slow Query Points Improve Local-SGD for Stochastic Convex Optimization
by: Dahan, Tehila, et al.
Published: (2023)
by: Dahan, Tehila, et al.
Published: (2023)
When Is Compositional Reasoning Learnable from Verifiable Rewards?
by: Barzilai, Daniel, et al.
Published: (2026)
by: Barzilai, Daniel, et al.
Published: (2026)
Repetita Iuvant: Data Repetition Allows SGD to Learn High-Dimensional Multi-Index Functions
by: Arnaboldi, Luca, et al.
Published: (2024)
by: Arnaboldi, Luca, et al.
Published: (2024)
Deterministic Nonsmooth Nonconvex Optimization
by: Jordan, Michael I., et al.
Published: (2023)
by: Jordan, Michael I., et al.
Published: (2023)
Anon: Extrapolating Adaptivity Beyond SGD and Adam
by: Zhang, Yiheng, et al.
Published: (2026)
by: Zhang, Yiheng, et al.
Published: (2026)
Near-Optimal Streaming Heavy-Tailed Statistical Estimation with Clipped SGD
by: Das, Aniket, et al.
Published: (2024)
by: Das, Aniket, et al.
Published: (2024)
On the Stability of Nonlinear Dynamics in GD and SGD: Beyond Quadratic Potentials
by: Mulayoff, Rotem, et al.
Published: (2026)
by: Mulayoff, Rotem, et al.
Published: (2026)
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning
by: Vyas, Nikhil, et al.
Published: (2023)
by: Vyas, Nikhil, et al.
Published: (2023)
Controlling the Inductive Bias of Wide Neural Networks by Modifying the Kernel's Spectrum
by: Geifman, Amnon, et al.
Published: (2023)
by: Geifman, Amnon, et al.
Published: (2023)
Asymptotics of SGD in Sequence-Single Index Models and Single-Layer Attention Networks
by: Arnaboldi, Luca, et al.
Published: (2025)
by: Arnaboldi, Luca, et al.
Published: (2025)
Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
by: Ertan, Murat Bilgehan, et al.
Published: (2026)
High-dimensional Limit of SGD for Diagonal Linear Networks
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
by: Malaxechebarría, Begoña García, et al.
Published: (2026)
Correlating Cross-Iteration Noise for DP-SGD using Model Curvature
by: Gu, Xin, et al.
Published: (2025)
by: Gu, Xin, et al.
Published: (2025)
Computational-Statistical Gaps in Gaussian Single-Index Models
by: Damian, Alex, et al.
Published: (2024)
by: Damian, Alex, et al.
Published: (2024)
Enhancing Stochastic Optimization for Statistical Efficiency Using ROOT-SGD with Diminishing Stepsize
by: Li, Chris Junchi
Published: (2024)
by: Li, Chris Junchi
Published: (2024)
Statistical Query Lower Bounds for Smoothed Agnostic Learning
by: Diakonikolas, Ilias, et al.
Published: (2026)
by: Diakonikolas, Ilias, et al.
Published: (2026)
A Noise Sensitivity Exponent Controls Large Statistical-to-Computational Gaps in Single- and Multi-Index Models
by: Defilippis, Leonardo, et al.
Published: (2026)
by: Defilippis, Leonardo, et al.
Published: (2026)
Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
by: Kovačević, Filip, et al.
Published: (2026)
by: Kovačević, Filip, et al.
Published: (2026)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Statistical-Computational Trade-offs in Learning Multi-Index Models via Harmonic Analysis
by: Latourelle-Vigeant, Hugo, et al.
Published: (2026)
by: Latourelle-Vigeant, Hugo, et al.
Published: (2026)
An open dataset of neural networks for hypernetwork research
by: Kurtenbach, David, et al.
Published: (2025)
by: Kurtenbach, David, et al.
Published: (2025)
Beyond SGD, Without SVD: Proximal Subspace Iteration LoRA with Diagonal Fractional K-FAC
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
by: Almansoori, Abdulla Jasem, et al.
Published: (2026)
Similar Items
-
When Models Don't Collapse: On the Consistency of Iterative MLE
by: Barzilai, Daniel, et al.
Published: (2025) -
Generalization in Kernel Regression Under Realistic Assumptions
by: Barzilai, Daniel, et al.
Published: (2023) -
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
by: Barzilai, Daniel, et al.
Published: (2025) -
Simple Relative Deviation Bounds for Covariance and Gram Matrices
by: Barzilai, Daniel, et al.
Published: (2024) -
Are Convex Optimization Curves Convex?
by: Barzilai, Guy, et al.
Published: (2025)