Expanding Expressivity in Transformer Models with MöbiusAttention
Fuente:
arXiv
Saved in:
| Main Authors: | Halacheva, Anna-Maria, Nayyeri, Mojtaba, Staab, Steffen |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Geometric Structural Knowledge Graph Foundation Model
by: Xin, Ling, et al.
Published: (2025)
by: Xin, Ling, et al.
Published: (2025)
Modeling Relational Patterns for Logical Query Answering over Knowledge Graphs
by: He, Yunjie, et al.
Published: (2023)
by: He, Yunjie, et al.
Published: (2023)
Generating $SROI^-$ Ontologies via Knowledge Graph Query Embedding Learning
by: He, Yunjie, et al.
Published: (2024)
by: He, Yunjie, et al.
Published: (2024)
Mathematical Reasoning for Unmanned Aerial Vehicles: A RAG-Based Approach for Complex Arithmetic Reasoning
by: Azarafza, Mehdi, et al.
Published: (2025)
by: Azarafza, Mehdi, et al.
Published: (2025)
Full-History Graphs with Edge-Type Decoupled Networks for Temporal Reasoning
by: Mohammed, Osama, et al.
Published: (2025)
by: Mohammed, Osama, et al.
Published: (2025)
Hybrid Reasoning Based on Large Language Models for Autonomous Car Driving
by: Azarafza, Mehdi, et al.
Published: (2024)
by: Azarafza, Mehdi, et al.
Published: (2024)
LMT-Net: Lane Model Transformer Network for Automated HD Mapping from Sparse Vehicle Observations
by: Mink, Michael, et al.
Published: (2024)
by: Mink, Michael, et al.
Published: (2024)
Towards Foundation Model on Temporal Knowledge Graph Reasoning
by: Pan, Jiaxin, et al.
Published: (2025)
by: Pan, Jiaxin, et al.
Published: (2025)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
by: De Schouwer, Jonas, et al.
Published: (2026)
by: De Schouwer, Jonas, et al.
Published: (2026)
SEMMA: A Semantic Aware Knowledge Graph Foundation Model
by: Arun, Arvindh, et al.
Published: (2025)
by: Arun, Arvindh, et al.
Published: (2025)
Leveraging Graph Structure in Seq2Seq Models for Knowledge Graph Link Prediction
by: Phuc, Luu Huu, et al.
Published: (2026)
by: Phuc, Luu Huu, et al.
Published: (2026)
Predictive Multiplicity of Knowledge Graph Embeddings in Link Prediction
by: Zhu, Yuqicheng, et al.
Published: (2024)
by: Zhu, Yuqicheng, et al.
Published: (2024)
More Expressive Attention with Negative Weights
by: Lv, Ang, et al.
Published: (2024)
by: Lv, Ang, et al.
Published: (2024)
Towards Foundation Models for Relational Databases with Language Models and Graph Neural Networks
by: Wu, Jingcheng, et al.
Published: (2026)
by: Wu, Jingcheng, et al.
Published: (2026)
Polynormer: Polynomial-Expressive Graph Transformer in Linear Time
by: Deng, Chenhui, et al.
Published: (2024)
by: Deng, Chenhui, et al.
Published: (2024)
Approximating Probabilistic Inference in Statistical EL with Knowledge Graph Embeddings
by: Zhu, Yuqicheng, et al.
Published: (2024)
by: Zhu, Yuqicheng, et al.
Published: (2024)
Pay Attention to the Triggers: Constructing Backdoors That Survive Distillation
by: De Muri, Giovanni, et al.
Published: (2025)
by: De Muri, Giovanni, et al.
Published: (2025)
Taipan: Efficient and Expressive State Space Language Models with Selective Attention
by: Van Nguyen, Chien, et al.
Published: (2024)
by: Van Nguyen, Chien, et al.
Published: (2024)
EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers
by: Liao, Yi-Lun, et al.
Published: (2026)
by: Liao, Yi-Lun, et al.
Published: (2026)
Is Complex Query Answering Really Complex?
by: Gregucci, Cosimo, et al.
Published: (2024)
by: Gregucci, Cosimo, et al.
Published: (2024)
On the Expressive Power of Transformers for Maxout Networks and Continuous Piecewise Linear Functions
by: Gu, Linyan, et al.
Published: (2026)
by: Gu, Linyan, et al.
Published: (2026)
Probabilistic Regular Tree Priors for Scientific Symbolic Reasoning
by: Schneider, Tim, et al.
Published: (2023)
by: Schneider, Tim, et al.
Published: (2023)
Emergence of Frontier Superposition: Möbius attractor and Cascade Supervision
by: Gu, Hongyu, et al.
Published: (2026)
by: Gu, Hongyu, et al.
Published: (2026)
Positional Attention: Expressivity and Learnability of Algorithmic Computation
by: de Luca, Artur Back, et al.
Published: (2024)
by: de Luca, Artur Back, et al.
Published: (2024)
Watermark Stealing in Large Language Models
by: Jovanović, Nikola, et al.
Published: (2024)
by: Jovanović, Nikola, et al.
Published: (2024)
Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers
by: Li, Xiaoyu, et al.
Published: (2024)
by: Li, Xiaoyu, et al.
Published: (2024)
Are Expressive Models Truly Necessary for Offline RL?
by: Wang, Guan, et al.
Published: (2024)
by: Wang, Guan, et al.
Published: (2024)
How Expressive are Knowledge Graph Foundation Models?
by: Huang, Xingyue, et al.
Published: (2025)
by: Huang, Xingyue, et al.
Published: (2025)
A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
by: Ye, Xiaowei, et al.
Published: (2026)
by: Ye, Xiaowei, et al.
Published: (2026)
Every Bit Counts: A Theoretical Study of Precision-Expressivity Tradeoffs in Quantized Transformers
by: Chakrabarti, Sayak, et al.
Published: (2026)
by: Chakrabarti, Sayak, et al.
Published: (2026)
GR-Agent: Adaptive Graph Reasoning Agent under Incomplete Knowledge
by: Zhou, Dongzhuoran, et al.
Published: (2025)
by: Zhou, Dongzhuoran, et al.
Published: (2025)
Yet Unnoticed in LSTM: Binary Tree Based Input Reordering, Weight Regularization, and Gate Nonlinearization
by: Moattari, Mojtaba
Published: (2025)
by: Moattari, Mojtaba
Published: (2025)
A Study on Variants of Conventional, Fuzzy, and Nullspace-Based Independence Criteria for Improving Supervised and Unsupervised Learning
by: Moattari, Mojtaba
Published: (2025)
by: Moattari, Mojtaba
Published: (2025)
NavMapFusion: Diffusion-based Fusion of Navigation Maps for Online Vectorized HD Map Construction
by: Monninger, Thomas, et al.
Published: (2025)
by: Monninger, Thomas, et al.
Published: (2025)
Remove Symmetries to Control Model Expressivity and Improve Optimization
by: Ziyin, Liu, et al.
Published: (2024)
by: Ziyin, Liu, et al.
Published: (2024)
Implicit Language Models are RNNs: Balancing Parallelization and Expressivity
by: Schöne, Mark, et al.
Published: (2025)
by: Schöne, Mark, et al.
Published: (2025)
CombAlign: Enhancing Model Expressiveness in Unsupervised Graph Alignment
by: Chen, Songyang, et al.
Published: (2024)
by: Chen, Songyang, et al.
Published: (2024)
Building Expressive and Tractable Probabilistic Generative Models: A Review
by: Sidheekh, Sahil, et al.
Published: (2024)
by: Sidheekh, Sahil, et al.
Published: (2024)
InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model
by: Wang, Youjin, et al.
Published: (2026)
by: Wang, Youjin, et al.
Published: (2026)
The Bayesian Geometry of Transformer Attention
by: Agarwal, Naman, et al.
Published: (2025)
by: Agarwal, Naman, et al.
Published: (2025)
Similar Items
-
Geometric Structural Knowledge Graph Foundation Model
by: Xin, Ling, et al.
Published: (2025) -
Modeling Relational Patterns for Logical Query Answering over Knowledge Graphs
by: He, Yunjie, et al.
Published: (2023) -
Generating $SROI^-$ Ontologies via Knowledge Graph Query Embedding Learning
by: He, Yunjie, et al.
Published: (2024) -
Mathematical Reasoning for Unmanned Aerial Vehicles: A RAG-Based Approach for Complex Arithmetic Reasoning
by: Azarafza, Mehdi, et al.
Published: (2025) -
Full-History Graphs with Edge-Type Decoupled Networks for Temporal Reasoning
by: Mohammed, Osama, et al.
Published: (2025)