The Bayesian Geometry of Transformer Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Agarwal, Naman, Dalal, Siddhartha R., Misra, Vishal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
by: Dalal, Siddhartha, et al.
Published: (2024)
by: Dalal, Siddhartha, et al.
Published: (2024)
Exact Attention Sensitivity and the Geometry of Transformer Stability
by: Emadi, Seyed Morteza
Published: (2026)
by: Emadi, Seyed Morteza
Published: (2026)
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)
by: Lu, Jiecheng, et al.
Published: (2026)
LUMOS: Large User MOdels for User Behavior Prediction
by: Nigam, Dhruv, et al.
Published: (2025)
by: Nigam, Dhruv, et al.
Published: (2025)
Training neural networks faster with minimal tuning using pre-computed lists of hyperparameters for NAdamW
by: Medapati, Sourabh, et al.
Published: (2025)
by: Medapati, Sourabh, et al.
Published: (2025)
ARDDQN: Attention Recurrent Double Deep Q-Network for UAV Coverage Path Planning and Data Harvesting
by: Kumar, Praveen, et al.
Published: (2024)
by: Kumar, Praveen, et al.
Published: (2024)
Phase-Aware Deep Learning with Complex-Valued CNNs for Audio Signal Applications
by: Agrawal, Naman
Published: (2025)
by: Agrawal, Naman
Published: (2025)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
by: Koh, Minsu, et al.
Published: (2025)
by: Koh, Minsu, et al.
Published: (2025)
Reliable Policy Iteration: Performance Robustness Across Architecture and Environment Perturbations
by: Eshwar, S. R., et al.
Published: (2025)
by: Eshwar, S. R., et al.
Published: (2025)
Sample Complexity Analysis for Constrained Bilevel Reinforcement Learning
by: Saxena, Naman, et al.
Published: (2026)
by: Saxena, Naman, et al.
Published: (2026)
A comprehensive study of on-device NLP applications -- VQA, automated Form filling, Smart Replies for Linguistic Codeswitching
by: Goyal, Naman
Published: (2024)
by: Goyal, Naman
Published: (2024)
Kalman Bayesian Transformer
by: Jing, Haoming, et al.
Published: (2025)
by: Jing, Haoming, et al.
Published: (2025)
Molecular De Novo Design through Transformer-based Reinforcement Learning
by: Xu, Pengcheng, et al.
Published: (2023)
by: Xu, Pengcheng, et al.
Published: (2023)
ConceptLens: from Pixels to Understanding
by: Dalal, Abhilekha, et al.
Published: (2024)
by: Dalal, Abhilekha, et al.
Published: (2024)
JAF: Judge Agent Forest
by: Garg, Sahil, et al.
Published: (2026)
by: Garg, Sahil, et al.
Published: (2026)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
by: Kumar, Phani, et al.
Published: (2026)
by: Kumar, Phani, et al.
Published: (2026)
Causal Pre-training Under the Fairness Lens: An Empirical Study of TabPFN
by: Liu, Qinyi, et al.
Published: (2026)
by: Liu, Qinyi, et al.
Published: (2026)
Provable Length Generalization in Sequence Prediction via Spectral Filtering
by: Marsden, Annie, et al.
Published: (2024)
by: Marsden, Annie, et al.
Published: (2024)
Transformers can do Bayesian Clustering
by: Bhaskaran, Prajit, et al.
Published: (2025)
by: Bhaskaran, Prajit, et al.
Published: (2025)
XicorAttention: Time Series Transformer Using Attention with Nonlinear Correlation
by: Kimura, Daichi, et al.
Published: (2025)
by: Kimura, Daichi, et al.
Published: (2025)
Foundation Priors
by: Misra, Sanjog
Published: (2025)
by: Misra, Sanjog
Published: (2025)
Geometric Attention: A Regime-Explicit Operator Semantics for Transformer Attention
by: Freytes, Luis Rosario
Published: (2026)
by: Freytes, Luis Rosario
Published: (2026)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
by: Miyato, Takeru, et al.
Published: (2023)
by: Miyato, Takeru, et al.
Published: (2023)
Attention Schema-based Attention Control (ASAC): A Cognitive-Inspired Approach for Attention Management in Transformers
by: Saxena, Krati, et al.
Published: (2025)
by: Saxena, Krati, et al.
Published: (2025)
Geometry-Aware Attention Guidance for Diffusion Models via Modern Hopfield Dynamics
by: Kim, Kwanyoung
Published: (2026)
by: Kim, Kwanyoung
Published: (2026)
Context-Sensitive Abstractions for Reinforcement Learning with Parameterized Actions
by: Nayyar, Rashmeet Kaur, et al.
Published: (2025)
by: Nayyar, Rashmeet Kaur, et al.
Published: (2025)
QuFeX: Quantum feature extraction module for hybrid quantum-classical deep neural networks
by: Jain, Naman, et al.
Published: (2025)
by: Jain, Naman, et al.
Published: (2025)
Can Vibe Coding Beat Graduate CS Students? An LLM vs. Human Coding Tournament on Market-driven Strategic Planning
by: Danassis, Panayiotis, et al.
Published: (2025)
by: Danassis, Panayiotis, et al.
Published: (2025)
A Post-Processing-Based Fair Federated Learning Framework
by: Zhou, Yi, et al.
Published: (2025)
by: Zhou, Yi, et al.
Published: (2025)
Unveiling and Controlling Anomalous Attention Distribution in Transformers
by: Yan, Ruiqing, et al.
Published: (2024)
by: Yan, Ruiqing, et al.
Published: (2024)
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
Higher-Order Transformers With Kronecker-Structured Attention
by: Omranpour, Soroush, et al.
Published: (2024)
by: Omranpour, Soroush, et al.
Published: (2024)
Expanding Expressivity in Transformer Models with MöbiusAttention
by: Halacheva, Anna-Maria, et al.
Published: (2024)
by: Halacheva, Anna-Maria, et al.
Published: (2024)
Transolver is a Linear Transformer: Revisiting Physics-Attention through the Lens of Linear Attention
by: Hu, Wenjie, et al.
Published: (2025)
by: Hu, Wenjie, et al.
Published: (2025)
Estimating Ground Reaction Forces from Inertial Sensors
by: Song, Bowen, et al.
Published: (2023)
by: Song, Bowen, et al.
Published: (2023)
Gradient Boosting Reinforcement Learning
by: Fuhrer, Benjamin, et al.
Published: (2024)
by: Fuhrer, Benjamin, et al.
Published: (2024)
Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
by: Fernando, Jesseba, et al.
Published: (2026)
by: Fernando, Jesseba, et al.
Published: (2026)
Echo State Transformer: Attention Over Finite Memories
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
by: Bendi-Ouis, Yannis, et al.
Published: (2025)
Similar Items
-
Gradient Dynamics of Attention: How Cross-Entropy Sculpts Bayesian Manifolds
by: Agarwal, Naman, et al.
Published: (2025) -
Geometric Scaling of Bayesian Inference in LLMs
by: Agarwal, Naman, et al.
Published: (2025) -
Beyond the Black Box: A Statistical Model for LLM Reasoning and Inference
by: Dalal, Siddhartha, et al.
Published: (2024) -
Exact Attention Sensitivity and the Geometry of Transformer Stability
by: Emadi, Seyed Morteza
Published: (2026) -
ZeroS: Zero-Sum Linear Attention for Efficient Transformers
by: Lu, Jiecheng, et al.
Published: (2026)