Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Kovačević, Filip, Ji, Hong Chang, Wu, Denny, Soltanolkotabi, Mahdi, Mondelli, Marco |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025)
by: Kovačević, Filip, et al.
Published: (2025)
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025)
by: Rezaei, Parham, et al.
Published: (2025)
Learning to Recall with Transformers Beyond Orthogonal Embeddings
by: Vural, Nuri Mert, et al.
Published: (2026)
by: Vural, Nuri Mert, et al.
Published: (2026)
Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing
by: Zhang, Yihan, et al.
Published: (2023)
by: Zhang, Yihan, et al.
Published: (2023)
Gradient Descent Provably Solves Nonlinear Tomographic Reconstruction
by: Fridovich-Keil, Sara, et al.
Published: (2023)
by: Fridovich-Keil, Sara, et al.
Published: (2023)
Optimal Estimation in Orthogonally Invariant Generalized Linear Models: Spectral Initialization and Approximate Message Passing
by: Zhang, Yihan, et al.
Published: (2026)
by: Zhang, Yihan, et al.
Published: (2026)
The Rich and the Simple: On the Implicit Bias of Adam and SGD
by: Vasudeva, Bhavya, et al.
Published: (2025)
by: Vasudeva, Bhavya, et al.
Published: (2025)
Test-Time Training Provably Improves Transformers as In-context Learners
by: Gozeten, Halil Alperen, et al.
Published: (2025)
by: Gozeten, Halil Alperen, et al.
Published: (2025)
From One-Pass SGD to Data Reuse: Mini-Batch Scaling Laws in Sketched Linear Regression
by: Chen, Ziyan, et al.
Published: (2026)
by: Chen, Ziyan, et al.
Published: (2026)
Adapt and Diffuse: Sample-adaptive Reconstruction via Latent Diffusion Models
by: Fabian, Zalan, et al.
Published: (2023)
by: Fabian, Zalan, et al.
Published: (2023)
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training
by: Zhao, Shen-Yi, et al.
Published: (2020)
by: Zhao, Shen-Yi, et al.
Published: (2020)
Convergence of Riemannian Stochastic Gradient Descents: Varying Batch Sizes And Nonstandard Batch Forming
by: Wu, Hao
Published: (2026)
by: Wu, Hao
Published: (2026)
AutoSGD: Automatic Learning Rate Selection for Stochastic Gradient Descent
by: Surjanovic, Nikola, et al.
Published: (2025)
by: Surjanovic, Nikola, et al.
Published: (2025)
Mini-Batch Covariance, Diffusion Limits, and Oracle Complexity in Stochastic Gradient Descent: A Sampling-Design Perspective
by: Zantedeschi, Daniel, et al.
Published: (2026)
by: Zantedeschi, Daniel, et al.
Published: (2026)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
by: Rajaraman, Nived, et al.
Published: (2026)
by: Rajaraman, Nived, et al.
Published: (2026)
ORFit: One-Pass Learning via Bridging Orthogonal Gradient Descent and Recursive Least-Squares
by: Min, Youngjae, et al.
Published: (2022)
by: Min, Youngjae, et al.
Published: (2022)
Asymmetric Prompt Weighting for Reinforcement Learning with Verifiable Rewards
by: Heckel, Reinhard, et al.
Published: (2026)
by: Heckel, Reinhard, et al.
Published: (2026)
MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
by: Sepehri, Mohammad Shahab, et al.
Published: (2024)
Theoretical Insights into Overparameterized Models in Multi-Task and Replay-Based Continual Learning
by: Banayeeanzade, Amin, et al.
Published: (2024)
by: Banayeeanzade, Amin, et al.
Published: (2024)
Accelerating Single-Pass SGD for Generalized Linear Prediction
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Scaling Laws of SignSGD in Linear Regression: When Does It Outperform SGD?
by: Kim, Jihwan, et al.
Published: (2026)
by: Kim, Jihwan, et al.
Published: (2026)
The Sample Complexity of Gradient Descent in Stochastic Convex Optimization
by: Livni, Roi
Published: (2024)
by: Livni, Roi
Published: (2024)
On the Convergence of Wasserstein Gradient Descent for Sampling
by: Ta, Van Chien, et al.
Published: (2026)
by: Ta, Van Chien, et al.
Published: (2026)
Gradient Descent Efficiency Index
by: Dhingra, Aviral
Published: (2024)
by: Dhingra, Aviral
Published: (2024)
Optimal Growth Schedules for Batch Size and Learning Rate in SGD that Reduce SFO Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Neural Collapse Beyond the Unconstrained Features Model: Landscape, Dynamics, and Generalization in the Mean-Field Regime
by: Wu, Diyuan, et al.
Published: (2025)
by: Wu, Diyuan, et al.
Published: (2025)
MosaicMRI: A Diverse Dataset and Benchmark for Raw Musculoskeletal MRI
by: Arguello, Paula, et al.
Published: (2026)
by: Arguello, Paula, et al.
Published: (2026)
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent
by: Umeda, Hikaru, et al.
Published: (2024)
by: Umeda, Hikaru, et al.
Published: (2024)
DiracDiffusion: Denoising and Incremental Reconstruction with Assured Data-Consistency
by: Fabian, Zalan, et al.
Published: (2023)
by: Fabian, Zalan, et al.
Published: (2023)
Training Dynamics of Softmax Self-Attention: Fast Global Convergence via Preconditioning
by: Goel, Gautam, et al.
Published: (2026)
by: Goel, Gautam, et al.
Published: (2026)
Emergence and Evolution of Interpretable Concepts in Diffusion Models
by: Tinaz, Berk, et al.
Published: (2025)
by: Tinaz, Berk, et al.
Published: (2025)
Learning a Single Index Model from Anisotropic Data with vanilla Stochastic Gradient Descent
by: Braun, Guillaume, et al.
Published: (2025)
by: Braun, Guillaume, et al.
Published: (2025)
On the Utility of Equal Batch Sizes for Inference in Stochastic Gradient Descent
by: Singh, Rahul, et al.
Published: (2023)
by: Singh, Rahul, et al.
Published: (2023)
FoNE: Precise Single-Token Number Embeddings via Fourier Features
by: Zhou, Tianyi, et al.
Published: (2025)
by: Zhou, Tianyi, et al.
Published: (2025)
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
by: Kunstner, Frederik, et al.
Published: (2024)
by: Kunstner, Frederik, et al.
Published: (2024)
Why is parameter averaging beneficial in SGD? An objective smoothing perspective
by: Nitanda, Atsushi, et al.
Published: (2023)
by: Nitanda, Atsushi, et al.
Published: (2023)
On the Convergence of Stochastic Gradient Descent with Perturbed Forward-Backward Passes
by: Kong, Boao, et al.
Published: (2026)
by: Kong, Boao, et al.
Published: (2026)
Adaptive Batch Size and Learning Rate Scheduler for Stochastic Gradient Descent Based on Minimization of Stochastic First-order Oracle Complexity
by: Umeda, Hikaru, et al.
Published: (2025)
by: Umeda, Hikaru, et al.
Published: (2025)
Learning quadratic neural networks in high dimensions: SGD dynamics and scaling laws
by: Arous, Gérard Ben, et al.
Published: (2025)
by: Arous, Gérard Ben, et al.
Published: (2025)
Faster Convergence of Riemannian Stochastic Gradient Descent with Increasing Batch Size
by: Oowada, Kanata, et al.
Published: (2025)
by: Oowada, Kanata, et al.
Published: (2025)
Similar Items
-
Spectral Estimators for Multi-Index Models: Precise Asymptotics and Optimal Weak Recovery
by: Kovačević, Filip, et al.
Published: (2025) -
High-dimensional Analysis of Synthetic Data Selection
by: Rezaei, Parham, et al.
Published: (2025) -
Learning to Recall with Transformers Beyond Orthogonal Embeddings
by: Vural, Nuri Mert, et al.
Published: (2026) -
Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing
by: Zhang, Yihan, et al.
Published: (2023) -
Gradient Descent Provably Solves Nonlinear Tomographic Reconstruction
by: Fridovich-Keil, Sara, et al.
Published: (2023)