Optimal L2 Regularization in High-dimensional Continual Linear Regression
Fuente:
arXiv
Saved in:
| Main Authors: | Karpel, Gilad, Moroshko, Edward, Levinstein, Ran, Meir, Ron, Soudry, Daniel, Evron, Itay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025)
by: Levinstein, Ran, et al.
Published: (2025)
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025)
by: Tsipory, Matan, et al.
Published: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024)
by: Goldfarb, Daniel, et al.
Published: (2024)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Data-dependent and Oracle Bounds on Forgetting in Continual Learning
by: Friedman, Lior, et al.
Published: (2024)
by: Friedman, Lior, et al.
Published: (2024)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
Directional-Clamp PPO
by: Karpel, Gilad, et al.
Published: (2025)
by: Karpel, Gilad, et al.
Published: (2025)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Tensor-Parallelism with Partially Synchronized Activations
by: Lamprecht, Itay, et al.
Published: (2025)
by: Lamprecht, Itay, et al.
Published: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
by: Chmiel, Brian, et al.
Published: (2025)
by: Chmiel, Brian, et al.
Published: (2025)
Scaling FP8 training to trillion-token LLMs
by: Fishman, Maxim, et al.
Published: (2024)
by: Fishman, Maxim, et al.
Published: (2024)
Analysis of the Identifying Regulation with Adversarial Surrogates Algorithm
by: Teichner, Ron, et al.
Published: (2024)
by: Teichner, Ron, et al.
Published: (2024)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
by: Norman, Yotam, et al.
Published: (2025)
by: Norman, Yotam, et al.
Published: (2025)
Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression
by: Ioushua, Shahar Stein, et al.
Published: (2023)
by: Ioushua, Shahar Stein, et al.
Published: (2023)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
by: Haroush, Matan, et al.
Published: (2025)
by: Haroush, Matan, et al.
Published: (2025)
High-Dimensional Private Linear Regression with Optimal Rates
by: Bombari, Simone, et al.
Published: (2025)
by: Bombari, Simone, et al.
Published: (2025)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
by: Chmiel, Brian, et al.
Published: (2021)
by: Chmiel, Brian, et al.
Published: (2021)
Normalized Architectures are Natively 4-Bit
by: Fishman, Maxim, et al.
Published: (2026)
by: Fishman, Maxim, et al.
Published: (2026)
Workspace Optimization: How to Train Your Agent
by: Sarafian, Elad, et al.
Published: (2026)
by: Sarafian, Elad, et al.
Published: (2026)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
by: Levinstein, B. A., et al.
Published: (2023)
by: Levinstein, B. A., et al.
Published: (2023)
Joint auto-encoders: a flexible multi-task learning framework
by: Epstein, Baruch, et al.
Published: (2017)
by: Epstein, Baruch, et al.
Published: (2017)
Statistical curriculum learning: An elimination algorithm achieving an oracle risk
by: Cohen, Omer, et al.
Published: (2024)
by: Cohen, Omer, et al.
Published: (2024)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
by: Hoffer, Elad, et al.
Published: (2026)
by: Hoffer, Elad, et al.
Published: (2026)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
by: Wu, Jingfeng, et al.
Published: (2025)
by: Wu, Jingfeng, et al.
Published: (2025)
Discrete-time Competing-Risks Regression with or without Penalization
by: Meir, Tomer, et al.
Published: (2023)
by: Meir, Tomer, et al.
Published: (2023)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
by: Clarté, Lucas, et al.
Published: (2024)
by: Clarté, Lucas, et al.
Published: (2024)
Optimal Implicit Bias in Linear Regression
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
by: Varma, Kanumuri Nithin, et al.
Published: (2025)
A Graph Meta-Network for Learning on Kolmogorov-Arnold Networks
by: Bar-Shalom, Guy, et al.
Published: (2026)
by: Bar-Shalom, Guy, et al.
Published: (2026)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
High-dimensional Asymptotics of Generalization Performance in Continual Ridge Regression
by: Zhao, Yihan, et al.
Published: (2025)
by: Zhao, Yihan, et al.
Published: (2025)
Understanding Forgetting in Continual Learning with Linear Regression
by: Ding, Meng, et al.
Published: (2024)
by: Ding, Meng, et al.
Published: (2024)
Characterization of the Distortion-Perception Tradeoff for Finite Channels with Arbitrary Metrics
by: Freirich, Dror, et al.
Published: (2024)
by: Freirich, Dror, et al.
Published: (2024)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
by: Zloczower, Itay, et al.
Published: (2026)
by: Zloczower, Itay, et al.
Published: (2026)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
by: Agrawalla, Bhavya, et al.
Published: (2023)
by: Agrawalla, Bhavya, et al.
Published: (2023)
Explore to Generalize in Zero-Shot RL
by: Zisselman, Ev, et al.
Published: (2023)
by: Zisselman, Ev, et al.
Published: (2023)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
by: Ravi, Hrithik, et al.
Published: (2024)
by: Ravi, Hrithik, et al.
Published: (2024)
$L_1$-norm Regularized Indefinite Kernel Logistic Regression
by: Wang, Shaoxin, et al.
Published: (2025)
by: Wang, Shaoxin, et al.
Published: (2025)
Similar Items
-
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025) -
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025) -
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025) -
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024) -
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)