Optimal Rates in Continual Linear Regression via Increasing Regularization
Fuente:
arXiv
Saved in:
| Main Authors: | Levinstein, Ran, Attia, Amit, Schliserman, Matan, Sherman, Uri, Koren, Tomer, Soudry, Daniel, Evron, Itay |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
Optimal L2 Regularization in High-dimensional Continual Linear Regression
by: Karpel, Gilad, et al.
Published: (2026)
by: Karpel, Gilad, et al.
Published: (2026)
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025)
by: Tsipory, Matan, et al.
Published: (2025)
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Complexity of Vector-valued Prediction: From Linear Models to Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)
by: Schliserman, Matan, et al.
Published: (2024)
Multiclass Loss Geometry Matters for Generalization of Gradient Descent in Separable Classification
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Rate-Optimal Policy Optimization for Linear Markov Decision Processes
by: Sherman, Uri, et al.
Published: (2023)
by: Sherman, Uri, et al.
Published: (2023)
Flat Minima and Generalization: Insights from Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2025)
by: Schliserman, Matan, et al.
Published: (2025)
Learning Rate Annealing Improves Tuning Robustness in Stochastic Optimization
by: Attia, Amit, et al.
Published: (2025)
by: Attia, Amit, et al.
Published: (2025)
How Free is Parameter-Free Stochastic Optimization?
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
A General Reduction for High-Probability Analysis with General Light-Tailed Distributions
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
Convergence of Policy Mirror Descent Beyond Compatible Function Approximation
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Convergence and Sample Complexity of First-Order Methods for Agnostic Reinforcement Learning
by: Sherman, Uri, et al.
Published: (2025)
by: Sherman, Uri, et al.
Published: (2025)
Faster Stochastic Optimization with Arbitrary Delays via Asynchronous Mini-Batching
by: Attia, Amit, et al.
Published: (2024)
by: Attia, Amit, et al.
Published: (2024)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024)
by: Goldfarb, Daniel, et al.
Published: (2024)
The Hidden Cost of Approximation in Online Mirror Descent
by: Schlisselberg, Ofir, et al.
Published: (2025)
by: Schlisselberg, Ofir, et al.
Published: (2025)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
by: Haroush, Matan, et al.
Published: (2025)
by: Haroush, Matan, et al.
Published: (2025)
Regret Minimization and Convergence to Equilibria in General-sum Markov Games
by: Erez, Liad, et al.
Published: (2022)
by: Erez, Liad, et al.
Published: (2022)
Nearly Optimal Sample Complexity for Learning with Label Proportions
by: Busa-Fekete, Robert, et al.
Published: (2025)
by: Busa-Fekete, Robert, et al.
Published: (2025)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
Towards Fully Parameter-Free Stochastic Optimization: Grid Search with Self-Bounding Analysis
by: Zhao, Yuheng, et al.
Published: (2026)
by: Zhao, Yuheng, et al.
Published: (2026)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Tensor-Parallelism with Partially Synchronized Activations
by: Lamprecht, Itay, et al.
Published: (2025)
by: Lamprecht, Itay, et al.
Published: (2025)
From Contextual Combinatorial Semi-Bandits to Bandit List Classification: Improved Sample Complexity with Sparse Rewards
by: Erez, Liad, et al.
Published: (2025)
by: Erez, Liad, et al.
Published: (2025)
High-Dimensional Private Linear Regression with Optimal Rates
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
State Entropy Regularization for Robust Reinforcement Learning
by: Ashlag, Yonatan, et al.
Published: (2025)
by: Ashlag, Yonatan, et al.
Published: (2025)
Optimal Learning from Label Proportions with General Loss Functions
by: Applebaum, Lorne, et al.
Published: (2025)
by: Applebaum, Lorne, et al.
Published: (2025)
Fast Rates for Bandit PAC Multiclass Classification
by: Erez, Liad, et al.
Published: (2024)
by: Erez, Liad, et al.
Published: (2024)
When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets
by: Zeno, Chen, et al.
Published: (2025)
by: Zeno, Chen, et al.
Published: (2025)
Multiplicative Reweighting for Robust Neural Network Optimization
by: Bar, Noga, et al.
Published: (2021)
by: Bar, Noga, et al.
Published: (2021)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
by: Levinstein, B. A., et al.
Published: (2023)
by: Levinstein, B. A., et al.
Published: (2023)
Heterogeneous Treatment Effect in Time-to-Event Outcomes: Harnessing Censored Data with Recursively Imputed Trees
by: Meir, Tomer, et al.
Published: (2025)
by: Meir, Tomer, et al.
Published: (2025)
On Local Overfitting and Forgetting in Deep Neural Networks
by: Stern, Uri, et al.
Published: (2024)
by: Stern, Uri, et al.
Published: (2024)
Rapid Overfitting of Multi-Pass Stochastic Gradient Descent in Stochastic Convex Optimization
by: Vansover-Hager, Shira, et al.
Published: (2025)
by: Vansover-Hager, Shira, et al.
Published: (2025)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Optimal Implicit Bias in Linear Regression
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
Similar Items
-
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025) -
Fast Last-Iterate Convergence of SGD in the Smooth Interpolation Regime
by: Attia, Amit, et al.
Published: (2025) -
Optimal L2 Regularization in High-dimensional Continual Linear Regression
by: Karpel, Gilad, et al.
Published: (2026) -
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025) -
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization
by: Schliserman, Matan, et al.
Published: (2024)