Adversarial Testing as a Tool for Interpretability: Length-based Overfitting of Elementary Functions in Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Zavoral, Patrik, Variš, Dušan, Bojar, Ondřej |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026)
by: Zhang, Jiaming, et al.
Published: (2026)
Testing for Overfitting
by: Schmidt, James
Published: (2023)
by: Schmidt, James
Published: (2023)
On Difficulties of Attention Factorization through Shared Memory
by: Yorsh, Uladzislau, et al.
Published: (2024)
by: Yorsh, Uladzislau, et al.
Published: (2024)
On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training
by: Liu, Chen, et al.
Published: (2021)
by: Liu, Chen, et al.
Published: (2021)
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
by: Hao, Yifan, et al.
Published: (2024)
by: Hao, Yifan, et al.
Published: (2024)
Out-of-distribution Tests Reveal Compositionality in Chess Transformers
by: Mészáros, Anna, et al.
Published: (2025)
by: Mészáros, Anna, et al.
Published: (2025)
Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples Regularization
by: Lin, Runqi, et al.
Published: (2024)
by: Lin, Runqi, et al.
Published: (2024)
Investigating Test Overfitting on SWE-bench
by: Ahmed, Toufique, et al.
Published: (2025)
by: Ahmed, Toufique, et al.
Published: (2025)
Higher-Order Message Passing for Glycan Representation Learning
by: Joeres, Roman, et al.
Published: (2024)
by: Joeres, Roman, et al.
Published: (2024)
Trained Transformer Classifiers Generalize and Exhibit Benign Overfitting In-Context
by: Frei, Spencer, et al.
Published: (2024)
by: Frei, Spencer, et al.
Published: (2024)
Adversarial training of Keyword Spotting to Minimize TTS Data Overfitting
by: Park, Hyun Jin, et al.
Published: (2024)
by: Park, Hyun Jin, et al.
Published: (2024)
Adjusted Overfitting Regression
by: Wilson, Dylan
Published: (2024)
by: Wilson, Dylan
Published: (2024)
Overfitting In Contrastive Learning?
by: Rabin, Zachary, et al.
Published: (2024)
by: Rabin, Zachary, et al.
Published: (2024)
Alleviating Overfitting in Transformation-Interaction-Rational Symbolic Regression with Multi-Objective Optimization
by: de Franca, Fabricio Olivetti
Published: (2025)
by: de Franca, Fabricio Olivetti
Published: (2025)
Prompting LLMs: Length Control for Isometric Machine Translation
by: Javorský, Dávid, et al.
Published: (2025)
by: Javorský, Dávid, et al.
Published: (2025)
Harmful Overfitting in Sobolev Spaces
by: Karhadkar, Kedar, et al.
Published: (2026)
by: Karhadkar, Kedar, et al.
Published: (2026)
Countering Overfitting with Counterfactual Examples
by: Giorgi, Flavio, et al.
Published: (2025)
by: Giorgi, Flavio, et al.
Published: (2025)
PINNs Failure Modes are Overfitting
by: Andersen, Nigel T., et al.
Published: (2026)
by: Andersen, Nigel T., et al.
Published: (2026)
Understanding Generalization in Transformers: Error Bounds and Training Dynamics Under Benign and Harmful Overfitting
by: Zhang, Yingying, et al.
Published: (2025)
by: Zhang, Yingying, et al.
Published: (2025)
On the Clean Generalization and Robust Overfitting in Adversarial Training from Two Theoretical Views: Representation Complexity and Training Dynamics
by: Li, Binghui, et al.
Published: (2023)
by: Li, Binghui, et al.
Published: (2023)
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
by: Zhao, Mengnan, et al.
Published: (2026)
by: Zhao, Mengnan, et al.
Published: (2026)
Control of Overfitting with Physics
by: Kozyrev, Sergei V., et al.
Published: (2024)
by: Kozyrev, Sergei V., et al.
Published: (2024)
Benign Overfitting in Single-Head Attention
by: Magen, Roey, et al.
Published: (2024)
by: Magen, Roey, et al.
Published: (2024)
Relative Overfitting and Accept-Reject Framework
by: Liu, Yanxin, et al.
Published: (2025)
by: Liu, Yanxin, et al.
Published: (2025)
Interpretability-Guided Test-Time Adversarial Defense
by: Kulkarni, Akshay, et al.
Published: (2024)
by: Kulkarni, Akshay, et al.
Published: (2024)
Interpretability in Parameter Space: Minimizing Mechanistic Description Length with Attribution-based Parameter Decomposition
by: Braun, Dan, et al.
Published: (2025)
by: Braun, Dan, et al.
Published: (2025)
Preventing Catastrophic Overfitting in Fast Adversarial Training: A Bi-level Optimization Perspective
by: Wang, Zhaoxin, et al.
Published: (2024)
by: Wang, Zhaoxin, et al.
Published: (2024)
Adversarially Diversified Rehearsal Memory (ADRM): Mitigating Memory Overfitting Challenge in Continual Learning
by: Khan, Hikmat, et al.
Published: (2024)
by: Khan, Hikmat, et al.
Published: (2024)
Looped Transformers for Length Generalization
by: Fan, Ying, et al.
Published: (2024)
by: Fan, Ying, et al.
Published: (2024)
Benign Overfitting in Linear Classifiers with a Bias Term
by: Kondo, Yuta
Published: (2025)
by: Kondo, Yuta
Published: (2025)
Benign Overfitting with Quantum Kernels
by: Tomasi, Joachim, et al.
Published: (2025)
by: Tomasi, Joachim, et al.
Published: (2025)
Overfitting in Adaptive Robust Optimization
by: Zhu, Karl, et al.
Published: (2025)
by: Zhu, Karl, et al.
Published: (2025)
Benign Overfitting in Token Selection of Attention Mechanism
by: Sakamoto, Keitaro, et al.
Published: (2024)
by: Sakamoto, Keitaro, et al.
Published: (2024)
On Local Overfitting and Forgetting in Deep Neural Networks
by: Stern, Uri, et al.
Published: (2024)
by: Stern, Uri, et al.
Published: (2024)
Unveil Benign Overfitting for Transformer in Vision: Training Dynamics, Convergence, and Generalization
by: Jiang, Jiarui, et al.
Published: (2024)
by: Jiang, Jiarui, et al.
Published: (2024)
Causality is Key for Interpretability Claims to Generalise
by: Joshi, Shruti, et al.
Published: (2026)
by: Joshi, Shruti, et al.
Published: (2026)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
Robust Federated Learning under Adversarial Attacks via Loss-Based Client Clustering
by: Kritharakis, Emmanouil, et al.
Published: (2025)
by: Kritharakis, Emmanouil, et al.
Published: (2025)
Catastrophic Overfitting, Entropy Gap and Participation Ratio: A Noiseless $l^p$ Norm Solution for Fast Adversarial Training
by: Mehouachi, Fares B., et al.
Published: (2025)
by: Mehouachi, Fares B., et al.
Published: (2025)
Provable Weak-to-Strong Generalization via Benign Overfitting
by: Wu, David X., et al.
Published: (2024)
by: Wu, David X., et al.
Published: (2024)
Similar Items
-
Benign Overfitting in Adversarial Training for Vision Transformers
by: Zhang, Jiaming, et al.
Published: (2026) -
Testing for Overfitting
by: Schmidt, James
Published: (2023) -
On Difficulties of Attention Factorization through Shared Memory
by: Yorsh, Uladzislau, et al.
Published: (2024) -
On the Impact of Hard Adversarial Instances on Overfitting in Adversarial Training
by: Liu, Chen, et al.
Published: (2021) -
The Surprising Harmfulness of Benign Overfitting for Adversarial Robustness
by: Hao, Yifan, et al.
Published: (2024)