Optimal L2 Regularization in High-dimensional Continual Linear Regression
Fuente:
arXiv
Salvato in:
| Autori principali: | Karpel, Gilad, Moroshko, Edward, Levinstein, Ran, Meir, Ron, Soudry, Daniel, Evron, Itay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Optimal Rates in Continual Linear Regression via Increasing Regularization
di: Levinstein, Ran, et al.
Pubblicazione: (2025)
di: Levinstein, Ran, et al.
Pubblicazione: (2025)
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
di: Tsipory, Matan, et al.
Pubblicazione: (2025)
di: Tsipory, Matan, et al.
Pubblicazione: (2025)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
di: Evron, Itay, et al.
Pubblicazione: (2025)
di: Evron, Itay, et al.
Pubblicazione: (2025)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
di: Goldfarb, Daniel, et al.
Pubblicazione: (2024)
di: Goldfarb, Daniel, et al.
Pubblicazione: (2024)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
di: Harel, Itamar, et al.
Pubblicazione: (2024)
di: Harel, Itamar, et al.
Pubblicazione: (2024)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
di: Chmiel, Brian, et al.
Pubblicazione: (2022)
di: Chmiel, Brian, et al.
Pubblicazione: (2022)
Data-dependent and Oracle Bounds on Forgetting in Continual Learning
di: Friedman, Lior, et al.
Pubblicazione: (2024)
di: Friedman, Lior, et al.
Pubblicazione: (2024)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
di: Blumenfeld, Yaniv, et al.
Pubblicazione: (2024)
di: Blumenfeld, Yaniv, et al.
Pubblicazione: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
di: Kinderman, Edan, et al.
Pubblicazione: (2024)
di: Kinderman, Edan, et al.
Pubblicazione: (2024)
Directional-Clamp PPO
di: Karpel, Gilad, et al.
Pubblicazione: (2025)
di: Karpel, Gilad, et al.
Pubblicazione: (2025)
Block Sparse Flash Attention
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
di: Ohayon, Daniel, et al.
Pubblicazione: (2025)
Tensor-Parallelism with Partially Synchronized Activations
di: Lamprecht, Itay, et al.
Pubblicazione: (2025)
di: Lamprecht, Itay, et al.
Pubblicazione: (2025)
FP4 All the Way: Fully Quantized Training of LLMs
di: Chmiel, Brian, et al.
Pubblicazione: (2025)
di: Chmiel, Brian, et al.
Pubblicazione: (2025)
Scaling FP8 training to trillion-token LLMs
di: Fishman, Maxim, et al.
Pubblicazione: (2024)
di: Fishman, Maxim, et al.
Pubblicazione: (2024)
Analysis of the Identifying Regulation with Adversarial Surrogates Algorithm
di: Teichner, Ron, et al.
Pubblicazione: (2024)
di: Teichner, Ron, et al.
Pubblicazione: (2024)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
di: Norman, Yotam, et al.
Pubblicazione: (2025)
di: Norman, Yotam, et al.
Pubblicazione: (2025)
Batches Stabilize the Minimum Norm Risk in High Dimensional Overparameterized Linear Regression
di: Ioushua, Shahar Stein, et al.
Pubblicazione: (2023)
di: Ioushua, Shahar Stein, et al.
Pubblicazione: (2023)
PLUMAGE: Probabilistic Low rank Unbiased Min Variance Gradient Estimator for Efficient Large Model Training
di: Haroush, Matan, et al.
Pubblicazione: (2025)
di: Haroush, Matan, et al.
Pubblicazione: (2025)
High-Dimensional Private Linear Regression with Optimal Rates
di: Bombari, Simone, et al.
Pubblicazione: (2025)
di: Bombari, Simone, et al.
Pubblicazione: (2025)
Accurate Neural Training with 4-bit Matrix Multiplications at Standard Formats
di: Chmiel, Brian, et al.
Pubblicazione: (2021)
di: Chmiel, Brian, et al.
Pubblicazione: (2021)
Normalized Architectures are Natively 4-Bit
di: Fishman, Maxim, et al.
Pubblicazione: (2026)
di: Fishman, Maxim, et al.
Pubblicazione: (2026)
Workspace Optimization: How to Train Your Agent
di: Sarafian, Elad, et al.
Pubblicazione: (2026)
di: Sarafian, Elad, et al.
Pubblicazione: (2026)
Still No Lie Detector for Language Models: Probing Empirical and Conceptual Roadblocks
di: Levinstein, B. A., et al.
Pubblicazione: (2023)
di: Levinstein, B. A., et al.
Pubblicazione: (2023)
Joint auto-encoders: a flexible multi-task learning framework
di: Epstein, Baruch, et al.
Pubblicazione: (2017)
di: Epstein, Baruch, et al.
Pubblicazione: (2017)
Statistical curriculum learning: An elimination algorithm achieving an oracle risk
di: Cohen, Omer, et al.
Pubblicazione: (2024)
di: Cohen, Omer, et al.
Pubblicazione: (2024)
Retrieval from Within: An Intrinsic Capability of Attention-Based Models
di: Hoffer, Elad, et al.
Pubblicazione: (2026)
di: Hoffer, Elad, et al.
Pubblicazione: (2026)
Risk Comparisons in Linear Regression: Implicit Regularization Dominates Explicit Regularization
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
di: Wu, Jingfeng, et al.
Pubblicazione: (2025)
Discrete-time Competing-Risks Regression with or without Penalization
di: Meir, Tomer, et al.
Pubblicazione: (2023)
di: Meir, Tomer, et al.
Pubblicazione: (2023)
Analysis of Bootstrap and Subsampling in High-dimensional Regularized Regression
di: Clarté, Lucas, et al.
Pubblicazione: (2024)
di: Clarté, Lucas, et al.
Pubblicazione: (2024)
Optimal Implicit Bias in Linear Regression
di: Varma, Kanumuri Nithin, et al.
Pubblicazione: (2025)
di: Varma, Kanumuri Nithin, et al.
Pubblicazione: (2025)
A Graph Meta-Network for Learning on Kolmogorov-Arnold Networks
di: Bar-Shalom, Guy, et al.
Pubblicazione: (2026)
di: Bar-Shalom, Guy, et al.
Pubblicazione: (2026)
To Grok Grokking: Provable Grokking in Ridge Regression
di: Xu, Mingyue, et al.
Pubblicazione: (2026)
di: Xu, Mingyue, et al.
Pubblicazione: (2026)
High-dimensional Asymptotics of Generalization Performance in Continual Ridge Regression
di: Zhao, Yihan, et al.
Pubblicazione: (2025)
di: Zhao, Yihan, et al.
Pubblicazione: (2025)
Understanding Forgetting in Continual Learning with Linear Regression
di: Ding, Meng, et al.
Pubblicazione: (2024)
di: Ding, Meng, et al.
Pubblicazione: (2024)
Characterization of the Distortion-Perception Tradeoff for Finite Channels with Arbitrary Metrics
di: Freirich, Dror, et al.
Pubblicazione: (2024)
di: Freirich, Dror, et al.
Pubblicazione: (2024)
One Step to the Side: Why Defenses Against Malicious Finetuning Fail Under Adaptive Adversaries
di: Zloczower, Itay, et al.
Pubblicazione: (2026)
di: Zloczower, Itay, et al.
Pubblicazione: (2026)
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression
di: Agrawalla, Bhavya, et al.
Pubblicazione: (2023)
di: Agrawalla, Bhavya, et al.
Pubblicazione: (2023)
Explore to Generalize in Zero-Shot RL
di: Zisselman, Ev, et al.
Pubblicazione: (2023)
di: Zisselman, Ev, et al.
Pubblicazione: (2023)
The Implicit Bias of Gradient Descent on Separable Multiclass Data
di: Ravi, Hrithik, et al.
Pubblicazione: (2024)
di: Ravi, Hrithik, et al.
Pubblicazione: (2024)
$L_1$-norm Regularized Indefinite Kernel Logistic Regression
di: Wang, Shaoxin, et al.
Pubblicazione: (2025)
di: Wang, Shaoxin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Optimal Rates in Continual Linear Regression via Increasing Regularization
di: Levinstein, Ran, et al.
Pubblicazione: (2025) -
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
di: Tsipory, Matan, et al.
Pubblicazione: (2025) -
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
di: Evron, Itay, et al.
Pubblicazione: (2025) -
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
di: Goldfarb, Daniel, et al.
Pubblicazione: (2024) -
Provable Tempered Overfitting of Minimal Nets and Typical Nets
di: Harel, Itamar, et al.
Pubblicazione: (2024)