Provable Tempered Overfitting of Minimal Nets and Typical Nets
Fuente:
arXiv
Saved in:
| Main Authors: | Harel, Itamar, Hoza, William M., Vardi, Gal, Evron, Itay, Srebro, Nathan, Soudry, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024)
by: Buzaglo, Gon, et al.
Published: (2024)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023)
by: Zhou, Lijia, et al.
Published: (2023)
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)
by: Evron, Itay, et al.
Published: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024)
by: Smorodinsky, Guy, et al.
Published: (2024)
The Joint Effect of Task Similarity and Overparameterization on Catastrophic Forgetting -- An Analytical Model
by: Goldfarb, Daniel, et al.
Published: (2024)
by: Goldfarb, Daniel, et al.
Published: (2024)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
Foldable SuperNets: Scalable Merging of Transformers with Different Initializations and Tasks
by: Kinderman, Edan, et al.
Published: (2024)
by: Kinderman, Edan, et al.
Published: (2024)
Quantifying Overfitting along the Regularization Path for Two-Part-Code MDL in Supervised Classification
by: Zhu, Xiaohan, et al.
Published: (2025)
by: Zhu, Xiaohan, et al.
Published: (2025)
Are Greedy Task Orderings Better Than Random in Continual Linear Regression?
by: Tsipory, Matan, et al.
Published: (2025)
by: Tsipory, Matan, et al.
Published: (2025)
Optimal L2 Regularization in High-dimensional Continual Linear Regression
by: Karpel, Gilad, et al.
Published: (2026)
by: Karpel, Gilad, et al.
Published: (2026)
Overfitting and Generalizing with (PAC) Bayesian Prediction in Noisy Binary Classification
by: Zhu, Xiaohan, et al.
Published: (2026)
by: Zhu, Xiaohan, et al.
Published: (2026)
Learning to Think from Multiple Thinkers
by: Joshi, Nirmit, et al.
Published: (2026)
by: Joshi, Nirmit, et al.
Published: (2026)
Positive Distribution Shift as a Framework for Understanding Tractable Learning
by: Medvedev, Marko, et al.
Published: (2026)
by: Medvedev, Marko, et al.
Published: (2026)
Optimal Rates in Continual Linear Regression via Increasing Regularization
by: Levinstein, Ran, et al.
Published: (2025)
by: Levinstein, Ran, et al.
Published: (2025)
The Implicit Bias of Gradient Descent on Separable Data
by: Soudry, Daniel, et al.
Published: (2017)
by: Soudry, Daniel, et al.
Published: (2017)
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
Malign Overfitting: Interpolation Can Provably Preclude Invariance
by: Wald, Yoav, et al.
Published: (2022)
by: Wald, Yoav, et al.
Published: (2022)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
Provable Unlearning with Gradient Ascent on Two-Layer ReLU Neural Networks
by: Melamed, Odelia, et al.
Published: (2025)
by: Melamed, Odelia, et al.
Published: (2025)
A Theory of Learning with Autoregressive Chain of Thought
by: Joshi, Nirmit, et al.
Published: (2025)
by: Joshi, Nirmit, et al.
Published: (2025)
Towards Cheaper Inference in Deep Networks with Lower Bit-Width Accumulators
by: Blumenfeld, Yaniv, et al.
Published: (2024)
by: Blumenfeld, Yaniv, et al.
Published: (2024)
Weak-to-Strong Generalization Even in Random Feature Networks, Provably
by: Medvedev, Marko, et al.
Published: (2025)
by: Medvedev, Marko, et al.
Published: (2025)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Transformers are almost optimal metalearners for linear classification
by: Magen, Roey, et al.
Published: (2025)
by: Magen, Roey, et al.
Published: (2025)
Minimum Variance Unbiased N:M Sparsity for the Neural Gradients
by: Chmiel, Brian, et al.
Published: (2022)
by: Chmiel, Brian, et al.
Published: (2022)
Block Sparse Flash Attention
by: Ohayon, Daniel, et al.
Published: (2025)
by: Ohayon, Daniel, et al.
Published: (2025)
Tensor-Parallelism with Partially Synchronized Activations
by: Lamprecht, Itay, et al.
Published: (2025)
by: Lamprecht, Itay, et al.
Published: (2025)
Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step Sizes
by: Qiao, Dan, et al.
Published: (2024)
by: Qiao, Dan, et al.
Published: (2024)
When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural Nets
by: Zeno, Chen, et al.
Published: (2025)
by: Zeno, Chen, et al.
Published: (2025)
Provably tuning the ElasticNet across instances
by: Balcan, Maria-Florina, et al.
Published: (2022)
by: Balcan, Maria-Florina, et al.
Published: (2022)
Parallelly Tempered Generative Adversarial Nets: Toward Stabilized Gradients
by: Sohn, Jinwon, et al.
Published: (2024)
by: Sohn, Jinwon, et al.
Published: (2024)
Provable Weak-to-Strong Generalization via Benign Overfitting
by: Wu, David X., et al.
Published: (2024)
by: Wu, David X., et al.
Published: (2024)
Provable Generalization in Overparameterized Neural Nets
by: Dhingra, Aviral
Published: (2025)
by: Dhingra, Aviral
Published: (2025)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
From Tempered to Benign Overfitting in ReLU Neural Networks
by: Kornowski, Guy, et al.
Published: (2023)
by: Kornowski, Guy, et al.
Published: (2023)
Similar Items
-
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025) -
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024) -
How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
by: Buzaglo, Gon, et al.
Published: (2024) -
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023) -
From Continual Learning to SGD and Back: Better Rates for Continual Linear Models
by: Evron, Itay, et al.
Published: (2025)