Failures and Successes of Cross-Validation for Early-Stopped Gradient Descent
Fuente:
arXiv
Saved in:
| Main Authors: | Patil, Pratik, Wu, Yuchen, Tibshirani, Ryan J. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Ridge Regularization for Out-of-Distribution Prediction
by: Patil, Pratik, et al.
Published: (2024)
by: Patil, Pratik, et al.
Published: (2024)
Revisiting Optimism and Model Complexity in the Wake of Overparameterized Machine Learning
by: Patil, Pratik, et al.
Published: (2024)
by: Patil, Pratik, et al.
Published: (2024)
Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences
by: Aolaritei, Liviu, et al.
Published: (2025)
by: Aolaritei, Liviu, et al.
Published: (2025)
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
by: Yang, Yingzhen, et al.
Published: (2024)
by: Yang, Yingzhen, et al.
Published: (2024)
Gradient Equilibrium in Online Learning: Theory and Applications
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
by: Angelopoulos, Anastasios N., et al.
Published: (2025)
Random Matrix Theory of Early-Stopped Gradient Flow: A Transient BBP Scenario
by: Coeurdoux, Florentin, et al.
Published: (2026)
by: Coeurdoux, Florentin, et al.
Published: (2026)
Implicit Regularization Paths of Weighted Neural Representations
by: Du, Jin-Hong, et al.
Published: (2024)
by: Du, Jin-Hong, et al.
Published: (2024)
Asymptotically free sketched ridge ensembles: Risks, cross-validation, and tuning
by: Patil, Pratik, et al.
Published: (2023)
by: Patil, Pratik, et al.
Published: (2023)
Multivariate Trend Filtering for Lattice Data
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2021)
by: Sadhanala, Veeranjaneyulu, et al.
Published: (2021)
Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot Tuning
by: Dang, Hien, et al.
Published: (2026)
by: Dang, Hien, et al.
Published: (2026)
Cross-validation: what does it estimate and how well does it do it?
by: Bates, Stephen, et al.
Published: (2021)
by: Bates, Stephen, et al.
Published: (2021)
Convex SGD: Generalization Without Early Stopping
by: Hendrickx, Julien, et al.
Published: (2024)
by: Hendrickx, Julien, et al.
Published: (2024)
Sharp Risk Bounds for Early-Stopping in Gaussian Linear Regression
by: Wegel, Tobias, et al.
Published: (2025)
by: Wegel, Tobias, et al.
Published: (2025)
A Stein Gradient Descent Approach for Doubly Intractable Distributions
by: Lee, Heesang, et al.
Published: (2024)
by: Lee, Heesang, et al.
Published: (2024)
Finite-Particle Rates for Regularized Stein Variational Gradient Descent
by: He, Ye, et al.
Published: (2026)
by: He, Ye, et al.
Published: (2026)
Interactive Learning of Single-Index Models via Stochastic Gradient Descent
by: Rajaraman, Nived, et al.
Published: (2026)
by: Rajaraman, Nived, et al.
Published: (2026)
Cross-regularization: Adaptive Model Complexity through Validation Gradients
by: Brito, Carlos Stein
Published: (2025)
by: Brito, Carlos Stein
Published: (2025)
Basic Inequalities for First-Order Optimization with Applications to Statistical Risk Analysis
by: Paik, Seunghoon, et al.
Published: (2025)
by: Paik, Seunghoon, et al.
Published: (2025)
Early Stopping in Contextual Bandits and Inferences
by: Cui, Zihan
Published: (2025)
by: Cui, Zihan
Published: (2025)
On Regularization via Early Stopping for Least Squares Regression
by: Sonthalia, Rishi, et al.
Published: (2024)
by: Sonthalia, Rishi, et al.
Published: (2024)
Precise Asymptotics of Bagging Regularized M-estimators
by: Koriyama, Takuya, et al.
Published: (2024)
by: Koriyama, Takuya, et al.
Published: (2024)
Improved Finite-Particle Convergence Rates for Stein Variational Gradient Descent
by: Banerjee, Sayan, et al.
Published: (2024)
by: Banerjee, Sayan, et al.
Published: (2024)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022)
by: Frei, Spencer, et al.
Published: (2022)
Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks
by: Rangriz, Parsa
Published: (2025)
by: Rangriz, Parsa
Published: (2025)
Implicit Regularization for Tubal Tensor Factorizations via Gradient Descent
by: Karnik, Santhosh, et al.
Published: (2024)
by: Karnik, Santhosh, et al.
Published: (2024)
Learning Operators with Stochastic Gradient Descent in General Hilbert Spaces
by: Shi, Lei, et al.
Published: (2024)
by: Shi, Lei, et al.
Published: (2024)
Corrected generalized cross-validation for finite ensembles of penalized estimators
by: Bellec, Pierre C., et al.
Published: (2023)
by: Bellec, Pierre C., et al.
Published: (2023)
Learning Operators by Regularized Stochastic Gradient Descent with Operator-valued Kernels
by: Yang, Jia-Qi, et al.
Published: (2025)
by: Yang, Jia-Qi, et al.
Published: (2025)
Gradient Descent with Projection Finds Over-Parameterized Neural Networks for Learning Low-Degree Polynomials with Nearly Minimax Optimal Rate
by: Yang, Yingzhen, et al.
Published: (2026)
by: Yang, Yingzhen, et al.
Published: (2026)
Bootstrapping the Cross-Validation Estimate
by: Cai, Bryan, et al.
Published: (2023)
by: Cai, Bryan, et al.
Published: (2023)
Dropout Drops Double Descent
by: Yang, Tian-Le, et al.
Published: (2023)
by: Yang, Tian-Le, et al.
Published: (2023)
The Structure of Cross-Validation Error: Stability, Covariance, and Minimax Limits
by: Nachum, Ido, et al.
Published: (2025)
by: Nachum, Ido, et al.
Published: (2025)
Minimax Limits of k-Fold Cross-Validation via Majority
by: Nachum, Ido, et al.
Published: (2026)
by: Nachum, Ido, et al.
Published: (2026)
ROTI-GCV: Generalized Cross-Validation for right-ROTationally Invariant Data
by: Luo, Kevin, et al.
Published: (2024)
by: Luo, Kevin, et al.
Published: (2024)
Differentially Private Two-Stage Gradient Descent for Instrumental Variable Regression
by: Liang, Haodong, et al.
Published: (2025)
by: Liang, Haodong, et al.
Published: (2025)
High-probability Convergence Bounds for Nonlinear Stochastic Gradient Descent Under Heavy-tailed Noise
by: Armacki, Aleksandar, et al.
Published: (2023)
by: Armacki, Aleksandar, et al.
Published: (2023)
On Stopping Times of Power-one Sequential Tests: Tight Lower and Upper Bounds
by: Agrawal, Shubhada, et al.
Published: (2025)
by: Agrawal, Shubhada, et al.
Published: (2025)
When Does Model Collapse Occur in Structured Interactive Learning?
by: Wu, Yuchen, et al.
Published: (2026)
by: Wu, Yuchen, et al.
Published: (2026)
Optimal Convergence Analysis of DDPM for General Distributions
by: Jiao, Yuchen, et al.
Published: (2025)
by: Jiao, Yuchen, et al.
Published: (2025)
From Cross-Validation to SURE: Asymptotic Risk of Tuned Regularized Estimators
by: Adusumilli, Karun, et al.
Published: (2026)
by: Adusumilli, Karun, et al.
Published: (2026)
Similar Items
-
Optimal Ridge Regularization for Out-of-Distribution Prediction
by: Patil, Pratik, et al.
Published: (2024) -
Revisiting Optimism and Model Complexity in the Wake of Overparameterized Machine Learning
by: Patil, Pratik, et al.
Published: (2024) -
Stopping Rules for Stochastic Gradient Descent via Anytime-Valid Confidence Sequences
by: Aolaritei, Liviu, et al.
Published: (2025) -
Sharp Generalization for Nonparametric Regression in Interpolation Space by Over-Parameterized Neural Networks Trained with Preconditioned Gradient Descent and Early Stopping
by: Yang, Yingzhen, et al.
Published: (2024) -
Gradient Equilibrium in Online Learning: Theory and Applications
by: Angelopoulos, Anastasios N., et al.
Published: (2025)