Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic Regression
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jingfeng, Bartlett, Peter, Telgarsky, Matus, Yu, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024)
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025)
On Achieving Optimal Adversarial Test Error
von: Li, Justin D., et al.
Veröffentlicht: (2023)
von: Li, Justin D., et al.
Veröffentlicht: (2023)
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)
Fast Robust Kernel Regression through Sign Gradient Descent with Early Stopping
von: Allerbo, Oskar
Veröffentlicht: (2023)
von: Allerbo, Oskar
Veröffentlicht: (2023)
Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast Optimization
von: Cai, Yuhang, et al.
Veröffentlicht: (2024)
von: Cai, Yuhang, et al.
Veröffentlicht: (2024)
Dual Space Preconditioning for Gradient Descent in the Overparameterized Regime
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
von: Ghane, Reza, et al.
Veröffentlicht: (2026)
Failures and Successes of Cross-Validation for Early-Stopped Gradient Descent
von: Patil, Pratik, et al.
Veröffentlicht: (2024)
von: Patil, Pratik, et al.
Veröffentlicht: (2024)
In-Context Learning of a Linear Transformer Block: Benefits of the MLP Component and One-Step GD Initialization
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2024)
Improved Scaling Laws in Linear Regression via Data Reuse
von: Lin, Licong, et al.
Veröffentlicht: (2025)
von: Lin, Licong, et al.
Veröffentlicht: (2025)
Transformers, parallel computation, and logarithmic depth
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
One-layer transformers fail to solve the induction heads task
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
von: Sanford, Clayton, et al.
Veröffentlicht: (2024)
GRADSTOP: Early Stopping of Gradient Descent via Posterior Sampling
von: Jamshidi, Arash, et al.
Veröffentlicht: (2025)
von: Jamshidi, Arash, et al.
Veröffentlicht: (2025)
Estimation of Toeplitz Covariance Matrices using Overparameterized Gradient Descent
von: Busbib, Daniel, et al.
Veröffentlicht: (2025)
von: Busbib, Daniel, et al.
Veröffentlicht: (2025)
Exponential Convergence of (Stochastic) Gradient Descent for Separable Logistic Regression
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
von: Kale, Sacchit, et al.
Veröffentlicht: (2026)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025)
Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models
von: Favero, Alessandro, et al.
Veröffentlicht: (2025)
von: Favero, Alessandro, et al.
Veröffentlicht: (2025)
Tight Bounds for Logistic Regression with Large Stepsize Gradient Descent in Low Dimension
von: Crawshaw, Michael, et al.
Veröffentlicht: (2026)
von: Crawshaw, Michael, et al.
Veröffentlicht: (2026)
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyang, et al.
Veröffentlicht: (2026)
Characterizing Dynamical Stability of Stochastic Gradient Descent in Overparameterized Learning
von: Chemnitz, Dennis, et al.
Veröffentlicht: (2024)
von: Chemnitz, Dennis, et al.
Veröffentlicht: (2024)
Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural Networks
von: Peleg, Amit, et al.
Veröffentlicht: (2024)
von: Peleg, Amit, et al.
Veröffentlicht: (2024)
From Logistic Regression to the Perceptron Algorithm: Exploring Gradient Descent with Large Step Sizes
von: Tyurin, Alexander
Veröffentlicht: (2024)
von: Tyurin, Alexander
Veröffentlicht: (2024)
Effectiveness of Distributed Gradient Descent with Local Steps for Overparameterized Models
von: Zhu, Heng, et al.
Veröffentlicht: (2024)
von: Zhu, Heng, et al.
Veröffentlicht: (2024)
Gradient Descent on Logistic Regression with Non-Separable Data and Large Step Sizes
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
von: Meng, Si Yi, et al.
Veröffentlicht: (2024)
Astral Space: Convex Analysis at Infinity
von: Dudík, Miroslav, et al.
Veröffentlicht: (2022)
von: Dudík, Miroslav, et al.
Veröffentlicht: (2022)
Stop Walking in Circles! Bailing Out Early in Projected Gradient Descent
von: Doldo, Philip, et al.
Veröffentlicht: (2025)
von: Doldo, Philip, et al.
Veröffentlicht: (2025)
Gradient Descent on Logistic Regression: Do Large Step-Sizes Work with Data on the Sphere?
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
von: Meng, Si Yi, et al.
Veröffentlicht: (2025)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
von: Yang, Yingzhen, et al.
Veröffentlicht: (2024)
Spectrum Extraction and Clipping for Implicitly Linear Layers
von: Boroojeny, Ali Ebrahimpour, et al.
Veröffentlicht: (2024)
von: Boroojeny, Ali Ebrahimpour, et al.
Veröffentlicht: (2024)
Preconditioned Gradient Descent for Overparameterized Nonconvex Burer--Monteiro Factorization with Global Optimality Certification
von: Zhang, Gavin, et al.
Veröffentlicht: (2022)
von: Zhang, Gavin, et al.
Veröffentlicht: (2022)
How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
von: Wu, Jingfeng, et al.
Veröffentlicht: (2023)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
von: Frei, Spencer, et al.
Veröffentlicht: (2022)
Transfer Learning of Linear Regression with Multiple Pretrained Models: Benefiting from More Pretrained Models via Overparameterization Debiasing
von: Boharon, Daniel, et al.
Veröffentlicht: (2026)
von: Boharon, Daniel, et al.
Veröffentlicht: (2026)
Stochastic Gradient Descent for Nonparametric Additive Regression
von: Chen, Xin, et al.
Veröffentlicht: (2024)
von: Chen, Xin, et al.
Veröffentlicht: (2024)
Scaling Laws in Linear Regression: Compute, Parameters, and Data
von: Lin, Licong, et al.
Veröffentlicht: (2024)
von: Lin, Licong, et al.
Veröffentlicht: (2024)
On the Benefits of Weight Normalization for Overparameterized Matrix Sensing
von: Wei, Yudong, et al.
Veröffentlicht: (2025)
von: Wei, Yudong, et al.
Veröffentlicht: (2025)
Overparameterized Multiple Linear Regression as Hyper-Curve Fitting
von: Atza, E., et al.
Veröffentlicht: (2024)
von: Atza, E., et al.
Veröffentlicht: (2024)
Learning Curves of Stochastic Gradient Descent in Kernel Regression
von: Zhang, Haihan, et al.
Veröffentlicht: (2025)
von: Zhang, Haihan, et al.
Veröffentlicht: (2025)
Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
von: Paik, Seunghoon, et al.
Veröffentlicht: (2025)
von: Paik, Seunghoon, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency
von: Wu, Jingfeng, et al.
Veröffentlicht: (2024) -
Large Stepsizes Accelerate Gradient Descent for Regularized Logistic Regression
von: Wu, Jingfeng, et al.
Veröffentlicht: (2025) -
Minimax Optimal Convergence of Gradient Descent in Logistic Regression via Large and Adaptive Stepsizes
von: Zhang, Ruiqi, et al.
Veröffentlicht: (2025) -
On Achieving Optimal Adversarial Test Error
von: Li, Justin D., et al.
Veröffentlicht: (2023) -
Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
von: Cai, Yuhang, et al.
Veröffentlicht: (2025)