Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
Fuente:
arXiv
Saved in:
| Main Authors: | Frei, Spencer, Vardi, Gal |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022)
by: Frei, Spencer, et al.
Published: (2022)
Benign Overfitting and the Geometry of the Ridge Regression Solution in Binary Classification
by: Tsigler, Alexander, et al.
Published: (2025)
by: Tsigler, Alexander, et al.
Published: (2025)
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024)
by: Medvedev, Marko, et al.
Published: (2024)
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Benign Overfitting in Linear Classifiers with a Bias Term
by: Kondo, Yuta
Published: (2025)
by: Kondo, Yuta
Published: (2025)
An Agnostic View on the Cost of Overfitting in (Kernel) Ridge Regression
by: Zhou, Lijia, et al.
Published: (2023)
by: Zhou, Lijia, et al.
Published: (2023)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Understanding Generalization in Transformers: Error Bounds and Training Dynamics Under Benign and Harmful Overfitting
by: Zhang, Yingying, et al.
Published: (2025)
by: Zhang, Yingying, et al.
Published: (2025)
Transformers are almost optimal metalearners for linear classification
by: Magen, Roey, et al.
Published: (2025)
by: Magen, Roey, et al.
Published: (2025)
Provable Tempered Overfitting of Minimal Nets and Typical Nets
by: Harel, Itamar, et al.
Published: (2024)
by: Harel, Itamar, et al.
Published: (2024)
Provable Weak-to-Strong Generalization via Benign Overfitting
by: Wu, David X., et al.
Published: (2024)
by: Wu, David X., et al.
Published: (2024)
Benign Overfitting in Out-of-Distribution Generalization of Linear Models
by: Tang, Shange, et al.
Published: (2024)
by: Tang, Shange, et al.
Published: (2024)
Benign Overfitting with Quantum Kernels
by: Tomasi, Joachim, et al.
Published: (2025)
by: Tomasi, Joachim, et al.
Published: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Provable Privacy Attacks on Trained Shallow Neural Networks
by: Smorodinsky, Guy, et al.
Published: (2024)
by: Smorodinsky, Guy, et al.
Published: (2024)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
by: Park, Junhyung, et al.
Published: (2024)
by: Park, Junhyung, et al.
Published: (2024)
The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
by: Gronich, Eitan, et al.
Published: (2026)
by: Gronich, Eitan, et al.
Published: (2026)
Rethinking Benign Overfitting in Two-Layer Neural Networks
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Universality of Benign Overfitting in Binary Linear Classification
by: Hashimoto, Ichiro, et al.
Published: (2025)
by: Hashimoto, Ichiro, et al.
Published: (2025)
Noisy Interpolation Learning with Shallow Univariate ReLU Networks
by: Joshi, Nirmit, et al.
Published: (2023)
by: Joshi, Nirmit, et al.
Published: (2023)
Implicit Regularization Towards Rank Minimization in ReLU Networks
by: Timor, Nadav, et al.
Published: (2022)
by: Timor, Nadav, et al.
Published: (2022)
To Grok Grokking: Provable Grokking in Ridge Regression
by: Xu, Mingyue, et al.
Published: (2026)
by: Xu, Mingyue, et al.
Published: (2026)
Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions
by: Lu, Weihao, et al.
Published: (2026)
by: Lu, Weihao, et al.
Published: (2026)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Transfer Learning for Benign Overfitting in High-Dimensional Linear Regression
by: Kim, Yeichan, et al.
Published: (2025)
by: Kim, Yeichan, et al.
Published: (2025)
Benign Overfitting in Time Series Linear Models with Over-Parameterization
by: Nakakita, Shogo, et al.
Published: (2022)
by: Nakakita, Shogo, et al.
Published: (2022)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
From Tempered to Benign Overfitting in ReLU Neural Networks
by: Kornowski, Guy, et al.
Published: (2023)
by: Kornowski, Guy, et al.
Published: (2023)
Malign Overfitting: Interpolation Can Provably Preclude Invariance
by: Wald, Yoav, et al.
Published: (2022)
by: Wald, Yoav, et al.
Published: (2022)
Benign Overfitting under Learning Rate Conditions for $α$ Sub-exponential Input
by: Okudo, Kota, et al.
Published: (2024)
by: Okudo, Kota, et al.
Published: (2024)
Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic Overfitting
by: Li, Jiping, et al.
Published: (2025)
by: Li, Jiping, et al.
Published: (2025)
Temperature is All You Need for Generalization in Langevin Dynamics and other Markov Processes
by: Harel, Itamar, et al.
Published: (2025)
by: Harel, Itamar, et al.
Published: (2025)
Understanding In-Context Learning of Linear Models in Transformers Through an Adversarial Lens
by: Anwar, Usman, et al.
Published: (2024)
by: Anwar, Usman, et al.
Published: (2024)
Initialization Matters: On the Benign Overfitting of Two-Layer ReLU CNN with Fully Trainable Layers
by: Shang, Shuning, et al.
Published: (2024)
by: Shang, Shuning, et al.
Published: (2024)
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
by: Tsilivis, Nikolaos, et al.
Published: (2024)
by: Tsilivis, Nikolaos, et al.
Published: (2024)
Provable Unlearning with Gradient Ascent on Two-Layer ReLU Neural Networks
by: Melamed, Odelia, et al.
Published: (2025)
by: Melamed, Odelia, et al.
Published: (2025)
On the Hardness of Learning Regular Expressions
by: Attias, Idan, et al.
Published: (2025)
by: Attias, Idan, et al.
Published: (2025)
Similar Items
-
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024) -
Benign Overfitting without Linearity: Neural Network Classifiers Trained by Gradient Descent for Noisy Linear Data
by: Frei, Spencer, et al.
Published: (2022) -
Benign Overfitting and the Geometry of the Ridge Regression Solution in Binary Classification
by: Tsigler, Alexander, et al.
Published: (2025) -
Overfitting Behaviour of Gaussian Kernel Ridgeless Regression: Varying Bandwidth or Dimensionality
by: Medvedev, Marko, et al.
Published: (2024) -
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026)