Exact Attention Sensitivity and the Geometry of Transformer Stability
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Emadi, Seyed Morteza |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
von: Emadi, Seyed Morteza
Veröffentlicht: (2026)
The Bayesian Geometry of Transformer Attention
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
von: Agarwal, Naman, et al.
Veröffentlicht: (2025)
Exact Linear Attention
von: Ou, Weinuo
Veröffentlicht: (2026)
von: Ou, Weinuo
Veröffentlicht: (2026)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
von: Koh, Minsu, et al.
Veröffentlicht: (2025)
von: Koh, Minsu, et al.
Veröffentlicht: (2025)
Parity, Sensitivity, and Transformers
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
von: Kozachinskiy, Alexander, et al.
Veröffentlicht: (2026)
Are Transformers More Robust? Towards Exact Robustness Verification for Transformers
von: Liao, Brian Hsuan-Cheng, et al.
Veröffentlicht: (2022)
von: Liao, Brian Hsuan-Cheng, et al.
Veröffentlicht: (2022)
Exact Dual Geometry of SOC-ICNN Value Functions
von: Liu, Kang, et al.
Veröffentlicht: (2026)
von: Liu, Kang, et al.
Veröffentlicht: (2026)
Geometry of Drifting MDPs with Path-Integral Stability Certificates
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zuyuan, et al.
Veröffentlicht: (2026)
N-vium: Mixture-of-Exits Transformer for Accelerated Exact Generation
von: Lorenc, Aleksander, et al.
Veröffentlicht: (2026)
von: Lorenc, Aleksander, et al.
Veröffentlicht: (2026)
Quantum Error Mitigation with Attention Graph Transformers for Burgers Equation Solvers on NISQ Hardware
von: Tousi, Seyed Mohamad Ali, et al.
Veröffentlicht: (2025)
von: Tousi, Seyed Mohamad Ali, et al.
Veröffentlicht: (2025)
On Exact Bit-level Reversible Transformers Without Changing Architectures
von: Zhang, Guoqiang, et al.
Veröffentlicht: (2024)
von: Zhang, Guoqiang, et al.
Veröffentlicht: (2024)
Representer Theorems for Metric and Preference Learning: Geometric Insights and Algorithms
von: Morteza, Peyman
Veröffentlicht: (2023)
von: Morteza, Peyman
Veröffentlicht: (2023)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
von: Deng, Yichuan, et al.
Veröffentlicht: (2024)
Vertex-Softmax: Tight Transformer Verification via Exact Softmax Optimization
von: Rezazadeh, Navid, et al.
Veröffentlicht: (2026)
von: Rezazadeh, Navid, et al.
Veröffentlicht: (2026)
Stability and Generalization in Looped Transformers
von: Labovich, Asher
Veröffentlicht: (2026)
von: Labovich, Asher
Veröffentlicht: (2026)
Noise Stability of Transformer Models
von: Haris, Themistoklis, et al.
Veröffentlicht: (2026)
von: Haris, Themistoklis, et al.
Veröffentlicht: (2026)
KANGURA: Kolmogorov-Arnold Network-Based Geometry-Aware Learning with Unified Representation Attention for 3D Modeling of Complex Structures
von: Shafie, Mohammad Reza, et al.
Veröffentlicht: (2025)
von: Shafie, Mohammad Reza, et al.
Veröffentlicht: (2025)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
von: Freytes, Luis Rosario
Veröffentlicht: (2026)
von: Freytes, Luis Rosario
Veröffentlicht: (2026)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
von: Kimura, Daichi, et al.
Veröffentlicht: (2025)
von: Kimura, Daichi, et al.
Veröffentlicht: (2025)
Less Is More -- On the Importance of Sparsification for Transformers and Graph Neural Networks for TSP
von: Lischka, Attila, et al.
Veröffentlicht: (2024)
von: Lischka, Attila, et al.
Veröffentlicht: (2024)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
von: Miyato, Takeru, et al.
Veröffentlicht: (2023)
von: Miyato, Takeru, et al.
Veröffentlicht: (2023)
Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
von: Saxena, Krati, et al.
Veröffentlicht: (2025)
von: Saxena, Krati, et al.
Veröffentlicht: (2025)
Geometry-Aware Attention Guidance for Diffusion Models via Modern Hopfield Dynamics
von: Kim, Kwanyoung
Veröffentlicht: (2026)
von: Kim, Kwanyoung
Veröffentlicht: (2026)
Shift of Pairwise Similarities for Data Clustering
von: Chehreghani, Morteza Haghir
Veröffentlicht: (2021)
von: Chehreghani, Morteza Haghir
Veröffentlicht: (2021)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
von: Yan, Ruiqing, et al.
Veröffentlicht: (2024)
Graph Convolutions Enrich the Self-Attention in Transformers!
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2023)
von: Choi, Jeongwhan, et al.
Veröffentlicht: (2023)
Higher-Order Transformers With Kronecker-Structured Attention
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
von: Omranpour, Soroush, et al.
Veröffentlicht: (2024)
Expanding Expressivity in Transformer Models with MöbiusAttention
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
von: Halacheva, Anna-Maria, et al.
Veröffentlicht: (2024)
Quantifying LLM Attention-Head Stability: Implications for Circuit Universality
von: Bali, Karan, et al.
Veröffentlicht: (2026)
von: Bali, Karan, et al.
Veröffentlicht: (2026)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
von: Hu, Wenjie, et al.
Veröffentlicht: (2025)
Measuring LLM Sensitivity in Transformer-based Tabular Data Synthesis
von: R, Maria F. Davila, et al.
Veröffentlicht: (2025)
von: R, Maria F. Davila, et al.
Veröffentlicht: (2025)
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
von: Fernando, Jesseba, et al.
Veröffentlicht: (2026)
von: Fernando, Jesseba, et al.
Veröffentlicht: (2026)
Decomposing Attention To Find Context-Sensitive Neurons
von: Gibson, Alex
Veröffentlicht: (2025)
von: Gibson, Alex
Veröffentlicht: (2025)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
von: Kumar, Phani, et al.
Veröffentlicht: (2026)
von: Kumar, Phani, et al.
Veröffentlicht: (2026)
Dispatch-Aware Ragged Attention for Pruned Vision Transformers
von: Abdellatif, Seifeldin, et al.
Veröffentlicht: (2026)
von: Abdellatif, Seifeldin, et al.
Veröffentlicht: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
von: Lu, Jiecheng, et al.
Veröffentlicht: (2026)
The Phasor Transformer: Resolving Attention Bottlenecks on the Unit Circle
von: Sigdel, Dibakar
Veröffentlicht: (2026)
von: Sigdel, Dibakar
Veröffentlicht: (2026)
Echo State Transformer: Attention Over Finite Memories
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2025)
von: Bendi-Ouis, Yannis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rank-Aware Spectral Bounds on Attention Logits for Stable Low-Precision Training
von: Emadi, Seyed Morteza
Veröffentlicht: (2026) -
The Causal Description Gap: Information-Theoretic Separations Across Pearl's Hierarchy
von: Emadi, Seyed Morteza
Veröffentlicht: (2026) -
The Critical Horizon: Inspection Design Principles for Multi-Stage Operations and Deep Reasoning
von: Emadi, Seyed Morteza
Veröffentlicht: (2026) -
The Bayesian Geometry of Transformer Attention
von: Agarwal, Naman, et al.
Veröffentlicht: (2025) -
Exact Linear Attention
von: Ou, Weinuo
Veröffentlicht: (2026)