Benign Overfitting in Token Selection of Attention Mechanism
Fuente:
arXiv
Saved in:
| Main Authors: | Sakamoto, Keitaro, Sato, Issei |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025)
by: Sakamoto, Keitaro, et al.
Published: (2025)
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
Benign Overfitting with Quantum Kernels
by: Tomasi, Joachim, et al.
Published: (2025)
by: Tomasi, Joachim, et al.
Published: (2025)
Rethinking Associative Memory Mechanism in Induction Head
by: Wang, Shuo, et al.
Published: (2024)
by: Wang, Shuo, et al.
Published: (2024)
Provable Weak-to-Strong Generalization via Benign Overfitting
by: Wu, David X., et al.
Published: (2024)
by: Wu, David X., et al.
Published: (2024)
Benign Overfitting in Out-of-Distribution Generalization of Linear Models
by: Tang, Shange, et al.
Published: (2024)
by: Tang, Shange, et al.
Published: (2024)
Rethinking Benign Overfitting in Two-Layer Neural Networks
by: Xu, Ruichen, et al.
Published: (2025)
by: Xu, Ruichen, et al.
Published: (2025)
Benign Overfitting in Linear Classifiers with a Bias Term
by: Kondo, Yuta
Published: (2025)
by: Kondo, Yuta
Published: (2025)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
Beyond Benign Overfitting in Nadaraya-Watson Interpolators
by: Barzilai, Daniel, et al.
Published: (2025)
by: Barzilai, Daniel, et al.
Published: (2025)
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Can Test-time Computation Mitigate Reproduction Bias in Neural Symbolic Regression?
by: Sato, Shun, et al.
Published: (2025)
by: Sato, Shun, et al.
Published: (2025)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
Benign Overfitting and the Geometry of the Ridge Regression Solution in Binary Classification
by: Tsigler, Alexander, et al.
Published: (2025)
by: Tsigler, Alexander, et al.
Published: (2025)
Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?
by: Kajitsuka, Tokio, et al.
Published: (2023)
by: Kajitsuka, Tokio, et al.
Published: (2023)
Universality of Benign Overfitting in Binary Linear Classification
by: Hashimoto, Ichiro, et al.
Published: (2025)
by: Hashimoto, Ichiro, et al.
Published: (2025)
Learning Curves and Benign Overfitting of Spectral Algorithms in Large Dimensions
by: Lu, Weihao, et al.
Published: (2026)
by: Lu, Weihao, et al.
Published: (2026)
A Classical View on Benign Overfitting: The Role of Sample Size
by: Park, Junhyung, et al.
Published: (2025)
by: Park, Junhyung, et al.
Published: (2025)
Transfer Learning for Benign Overfitting in High-Dimensional Linear Regression
by: Kim, Yeichan, et al.
Published: (2025)
by: Kim, Yeichan, et al.
Published: (2025)
Benign Overfitting in Time Series Linear Models with Over-Parameterization
by: Nakakita, Shogo, et al.
Published: (2022)
by: Nakakita, Shogo, et al.
Published: (2022)
Benign, Tempered, or Catastrophic: A Taxonomy of Overfitting
by: Mallinar, Neil, et al.
Published: (2022)
by: Mallinar, Neil, et al.
Published: (2022)
From Tempered to Benign Overfitting in ReLU Neural Networks
by: Kornowski, Guy, et al.
Published: (2023)
by: Kornowski, Guy, et al.
Published: (2023)
Multiplicative Logit Adjustment Approximates Neural-Collapse-Aware Decision Boundary Adjustment
by: Hasegawa, Naoya, et al.
Published: (2024)
by: Hasegawa, Naoya, et al.
Published: (2024)
On Expressive Power of Looped Transformers: Theoretical Analysis and Enhancement via Timestep Encoding
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
Understanding Linear Probing then Fine-tuning Language Models from NTK Perspective
by: Tomihari, Akiyoshi, et al.
Published: (2024)
by: Tomihari, Akiyoshi, et al.
Published: (2024)
Top-Down Bayesian Posterior Sampling for Sum-Product Networks
by: Yokoi, Soma, et al.
Published: (2024)
by: Yokoi, Soma, et al.
Published: (2024)
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation
by: Tomihari, Akiyoshi, et al.
Published: (2026)
by: Tomihari, Akiyoshi, et al.
Published: (2026)
Exploring Weight Balancing on Long-Tailed Recognition Problem
by: Hasegawa, Naoya, et al.
Published: (2023)
by: Hasegawa, Naoya, et al.
Published: (2023)
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
by: Park, Junhyung, et al.
Published: (2024)
by: Park, Junhyung, et al.
Published: (2024)
Understanding Generalization in Transformers: Error Bounds and Training Dynamics Under Benign and Harmful Overfitting
by: Zhang, Yingying, et al.
Published: (2025)
by: Zhang, Yingying, et al.
Published: (2025)
Benign Overfitting under Learning Rate Conditions for $α$ Sub-exponential Input
by: Okudo, Kota, et al.
Published: (2024)
by: Okudo, Kota, et al.
Published: (2024)
Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic Overfitting
by: Li, Jiping, et al.
Published: (2025)
by: Li, Jiping, et al.
Published: (2025)
Initialization Matters: On the Benign Overfitting of Two-Layer ReLU CNN with Fully Trainable Layers
by: Shang, Shuning, et al.
Published: (2024)
by: Shang, Shuning, et al.
Published: (2024)
On the Optimal Memorization Capacity of Transformers
by: Kajitsuka, Tokio, et al.
Published: (2024)
by: Kajitsuka, Tokio, et al.
Published: (2024)
Understanding Generalization in Physics Informed Models through Affine Variety Dimensions
by: Koshizuka, Takeshi, et al.
Published: (2025)
by: Koshizuka, Takeshi, et al.
Published: (2025)
Max-pooling Network Revisited: Analyzing the Role of Semantic Probability in Multiple Instance Learning for Hallucination Detection
by: Fujikawa, Shota, et al.
Published: (2026)
by: Fujikawa, Shota, et al.
Published: (2026)
To CoT or To Loop? A Formal Comparison Between Chain-of-Thought and Looped Transformers
by: Xu, Kevin, et al.
Published: (2025)
by: Xu, Kevin, et al.
Published: (2025)
Fix Initial Codes and Iteratively Refine Textual Directions Toward Safe Multi-Turn Code Correction
by: Tanaka, Yuto, et al.
Published: (2026)
by: Tanaka, Yuto, et al.
Published: (2026)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Similar Items
-
End-to-End Training Induces Information Bottleneck through Layer-Role Differentiation: A Comparative Analysis with Layer-wise Training
by: Sakamoto, Keitaro, et al.
Published: (2024) -
Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
by: Sakamoto, Keitaro, et al.
Published: (2025) -
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024) -
Benign Overfitting with Quantum Kernels
by: Tomasi, Joachim, et al.
Published: (2025) -
Rethinking Associative Memory Mechanism in Induction Head
by: Wang, Shuo, et al.
Published: (2024)