In Transformer We Trust? A Perspective on Transformer Architecture Failure Modes
Fuente:
arXiv
Salvato in:
| Autori principali: | Mondal, Trishit, Jagtap, Ameya D. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Hypersonic Flow Control: Generalized Deep Reinforcement Learning for Hypersonic Intake Unstart Control under Uncertainty
di: Mondal, Trishit, et al.
Pubblicazione: (2026)
di: Mondal, Trishit, et al.
Pubblicazione: (2026)
Shocks Under Control: Taming Transonic Compressible Flow over an RAE2822 Airfoil with Deep Reinforcement Learning
di: Mondal, Trishit, et al.
Pubblicazione: (2025)
di: Mondal, Trishit, et al.
Pubblicazione: (2025)
MARUT: An Exascale-Ready, GPU-Accelerated High-Order CFD Framework with AMR for High-Speed Flows and Finite-Rate Chemistry
di: Mondal, Trishit, et al.
Pubblicazione: (2026)
di: Mondal, Trishit, et al.
Pubblicazione: (2026)
An Approximation Theory Perspective on Machine Learning
di: Mhaskar, Hrushikesh N., et al.
Pubblicazione: (2025)
di: Mhaskar, Hrushikesh N., et al.
Pubblicazione: (2025)
Anant-Net: Breaking the Curse of Dimensionality with Scalable and Interpretable Neural Surrogate for High-Dimensional PDEs
di: Menon, Sidharth S., et al.
Pubblicazione: (2025)
di: Menon, Sidharth S., et al.
Pubblicazione: (2025)
FEKAN: Feature-Enriched Kolmogorov-Arnold Networks
di: Menon, Sidharth S., et al.
Pubblicazione: (2026)
di: Menon, Sidharth S., et al.
Pubblicazione: (2026)
BubbleOKAN: A Physics-Informed Interpretable Neural Operator for High-Frequency Bubble Dynamics
di: Zhang, Yunhao, et al.
Pubblicazione: (2025)
di: Zhang, Yunhao, et al.
Pubblicazione: (2025)
Understanding the Failure Modes of Transformers through the Lens of Graph Neural Networks
di: Lee, Hunjae
Pubblicazione: (2025)
di: Lee, Hunjae
Pubblicazione: (2025)
RiemannONets: Interpretable Neural Operators for Riemann Problems
di: Peyvan, Ahmad, et al.
Pubblicazione: (2024)
di: Peyvan, Ahmad, et al.
Pubblicazione: (2024)
Challenges and Advancements in Modeling Shock Fronts with Physics-Informed Neural Networks: A Review and Benchmarking Study
di: Abbasi, Jassem, et al.
Pubblicazione: (2025)
di: Abbasi, Jassem, et al.
Pubblicazione: (2025)
Multi-Narrow Transformation as a Single-Model Ensemble: Boundary Conditions, Mechanisms, and Failure Modes
di: Hasegawa, Tatsuhito, et al.
Pubblicazione: (2026)
di: Hasegawa, Tatsuhito, et al.
Pubblicazione: (2026)
Even Sparser Graph Transformers
di: Shirzad, Hamed, et al.
Pubblicazione: (2024)
di: Shirzad, Hamed, et al.
Pubblicazione: (2024)
A Theory for Compressibility of Graph Transformers for Transductive Learning
di: Shirzad, Hamed, et al.
Pubblicazione: (2024)
di: Shirzad, Hamed, et al.
Pubblicazione: (2024)
Trustformer: A Trusted Federated Transformer
di: Tadi, Ali Abbasi, et al.
Pubblicazione: (2025)
di: Tadi, Ali Abbasi, et al.
Pubblicazione: (2025)
Generalized Linear Mode Connectivity for Transformers
di: Theus, Alexander, et al.
Pubblicazione: (2025)
di: Theus, Alexander, et al.
Pubblicazione: (2025)
In Trust We Survive: Emergent Trust Learning
di: Chen, Qianpu, et al.
Pubblicazione: (2026)
di: Chen, Qianpu, et al.
Pubblicazione: (2026)
Hybrid Dual-Path Linear Transformations for Efficient Transformer Architectures
di: Khasia, Vladimer
Pubblicazione: (2026)
di: Khasia, Vladimer
Pubblicazione: (2026)
On Limitations of the Transformer Architecture
di: Peng, Binghui, et al.
Pubblicazione: (2024)
di: Peng, Binghui, et al.
Pubblicazione: (2024)
PINNs Failure Modes are Overfitting
di: Andersen, Nigel T., et al.
Pubblicazione: (2026)
di: Andersen, Nigel T., et al.
Pubblicazione: (2026)
Reducing the Transformer Architecture to a Minimum
di: Bermeitinger, Bernhard, et al.
Pubblicazione: (2024)
di: Bermeitinger, Bernhard, et al.
Pubblicazione: (2024)
Do We Need Transformers to Play FPS Video Games?
di: Batth, Karmanbir, et al.
Pubblicazione: (2025)
di: Batth, Karmanbir, et al.
Pubblicazione: (2025)
Expressivity of Transformers: A Tropical Geometry Perspective
di: Su, Ye, et al.
Pubblicazione: (2026)
di: Su, Ye, et al.
Pubblicazione: (2026)
A Constrained Optimization Perspective of Unrolled Transformers
di: Porras-Valenzuela, Javier, et al.
Pubblicazione: (2026)
di: Porras-Valenzuela, Javier, et al.
Pubblicazione: (2026)
Architecture Determines Observability of Transformers
di: Carmichael, Thomas
Pubblicazione: (2026)
di: Carmichael, Thomas
Pubblicazione: (2026)
Trust, but Verify: Peeling Low-Bit Transformer Networks for Training Monitoring
di: Eamaz, Arian, et al.
Pubblicazione: (2026)
di: Eamaz, Arian, et al.
Pubblicazione: (2026)
OT-Transformer: A Continuous-time Transformer Architecture with Optimal Transport Regularization
di: Kan, Kelvin, et al.
Pubblicazione: (2025)
di: Kan, Kelvin, et al.
Pubblicazione: (2025)
Separations in the Representational Capabilities of Transformers and Recurrent Architectures
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
di: Bhattamishra, Satwik, et al.
Pubblicazione: (2024)
Approximation Rate of the Transformer Architecture for Sequence Modeling
di: Jiang, Haotian, et al.
Pubblicazione: (2023)
di: Jiang, Haotian, et al.
Pubblicazione: (2023)
Understanding Transformer Architecture through Continuous Dynamics: A Partial Differential Equation Perspective
di: Zhang, Yukun, et al.
Pubblicazione: (2024)
di: Zhang, Yukun, et al.
Pubblicazione: (2024)
Are We Done with Object-Centric Learning?
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
di: Rubinstein, Alexander, et al.
Pubblicazione: (2025)
Finding Clustering Algorithms in the Transformer Architecture
di: Clarkson, Kenneth L., et al.
Pubblicazione: (2025)
di: Clarkson, Kenneth L., et al.
Pubblicazione: (2025)
Disentangling and Integrating Relational and Sensory Information in Transformer Architectures
di: Altabaa, Awni, et al.
Pubblicazione: (2024)
di: Altabaa, Awni, et al.
Pubblicazione: (2024)
Rethinking Graph Transformer Architecture Design for Node Classification
di: Zhou, Jiajun, et al.
Pubblicazione: (2024)
di: Zhou, Jiajun, et al.
Pubblicazione: (2024)
On the Universality of Transformer Architectures; How Much Attention Is Enough?
di: Abbasi, Amirreza, et al.
Pubblicazione: (2025)
di: Abbasi, Amirreza, et al.
Pubblicazione: (2025)
Anti Mode-Collapse in Mean-Field Transformer via Auxiliary Variables
di: Imaizumi, Masaaki, et al.
Pubblicazione: (2026)
di: Imaizumi, Masaaki, et al.
Pubblicazione: (2026)
A Simple Spectral Failure Mode for Graph Convolutional Networks
di: Priebe, Carey E., et al.
Pubblicazione: (2020)
di: Priebe, Carey E., et al.
Pubblicazione: (2020)
Born a Transformer -- Always a Transformer? On the Effect of Pretraining on Architectural Abilities
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
di: Jobanputra, Mayank, et al.
Pubblicazione: (2025)
Self-Attention as Distributional Projection: A Unified Interpretation of Transformer Architecture
di: Mehta, Nihal
Pubblicazione: (2025)
di: Mehta, Nihal
Pubblicazione: (2025)
A Transformer-based Autoregressive Decoder Architecture for Hierarchical Text Classification
di: Yousef, Younes, et al.
Pubblicazione: (2025)
di: Yousef, Younes, et al.
Pubblicazione: (2025)
Search-contempt: a hybrid MCTS algorithm for training AlphaZero-like engines with better computational efficiency
di: Joshi, Ameya
Pubblicazione: (2025)
di: Joshi, Ameya
Pubblicazione: (2025)
Documenti analoghi
-
Hypersonic Flow Control: Generalized Deep Reinforcement Learning for Hypersonic Intake Unstart Control under Uncertainty
di: Mondal, Trishit, et al.
Pubblicazione: (2026) -
Shocks Under Control: Taming Transonic Compressible Flow over an RAE2822 Airfoil with Deep Reinforcement Learning
di: Mondal, Trishit, et al.
Pubblicazione: (2025) -
MARUT: An Exascale-Ready, GPU-Accelerated High-Order CFD Framework with AMR for High-Speed Flows and Finite-Rate Chemistry
di: Mondal, Trishit, et al.
Pubblicazione: (2026) -
An Approximation Theory Perspective on Machine Learning
di: Mhaskar, Hrushikesh N., et al.
Pubblicazione: (2025) -
Anant-Net: Breaking the Curse of Dimensionality with Scalable and Interpretable Neural Surrogate for High-Dimensional PDEs
di: Menon, Sidharth S., et al.
Pubblicazione: (2025)