Saved in:
| Main Author: | Lin, Abigail |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.05483 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025)
by: Lin, Zhixuan, et al.
Published: (2025)
A Simple yet Effective DDG Predictor is An Unsupervised Antibody Optimizer and Explainer
by: Wu, Lirong, et al.
Published: (2025)
by: Wu, Lirong, et al.
Published: (2025)
Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
by: Gupta, Abhijit
Published: (2026)
by: Gupta, Abhijit
Published: (2026)
Crossfusor: A Cross-Attention Transformer Enhanced Conditional Diffusion Model for Car-Following Trajectory Prediction
by: You, Junwei, et al.
Published: (2024)
by: You, Junwei, et al.
Published: (2024)
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
by: Bu, Rui, et al.
Published: (2025)
by: Bu, Rui, et al.
Published: (2025)
Drug Synergy Prediction via Residual Graph Isomorphism Networks and Attention Mechanisms
by: Song, Jiyan, et al.
Published: (2026)
by: Song, Jiyan, et al.
Published: (2026)
JanusDDG: A Thermodynamics-Compliant Model for Sequence-Based Protein Stability via Two-Fronts Multi-Head Attention
by: Barducci, Guido, et al.
Published: (2025)
by: Barducci, Guido, et al.
Published: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
by: El, Batu, et al.
Published: (2025)
by: El, Batu, et al.
Published: (2025)
Graph Convolutions Enrich the Self-Attention in Transformers!
by: Choi, Jeongwhan, et al.
Published: (2023)
by: Choi, Jeongwhan, et al.
Published: (2023)
Revealing the Attention Floating Mechanism in Masked Diffusion Models
by: Dai, Xin, et al.
Published: (2026)
by: Dai, Xin, et al.
Published: (2026)
NoiseFormer -- Noise Diffused Symmetric Attention Transformer
by: Kumar, Phani, et al.
Published: (2026)
by: Kumar, Phani, et al.
Published: (2026)
Supercharging Graph Transformers with Advective Diffusion
by: Wu, Qitian, et al.
Published: (2023)
by: Wu, Qitian, et al.
Published: (2023)
A Neural Network Architecture Based on Attention Gate Mechanism for 3D Magnetotelluric Forward Modeling
by: Zhong, Xin, et al.
Published: (2025)
by: Zhong, Xin, et al.
Published: (2025)
BiTA: Bidirectional Gated Recurrent Unit-Transformer Aggregator in a Temporal Graph Network Framework for Alert Prediction in Computer Networks
by: Nayeri, Zahra Makki, et al.
Published: (2026)
by: Nayeri, Zahra Makki, et al.
Published: (2026)
Enhanced Graph Transformer with Serialized Graph Tokens
by: Wang, Ruixiang, et al.
Published: (2026)
by: Wang, Ruixiang, et al.
Published: (2026)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
by: Zhou, Jingbo, et al.
Published: (2026)
by: Zhou, Jingbo, et al.
Published: (2026)
YZS-model: A Predictive Model for Organic Drug Solubility Based on Graph Convolutional Networks and Transformer-Attention
by: Wang, Chenxu, et al.
Published: (2024)
by: Wang, Chenxu, et al.
Published: (2024)
Graph Diffusion Transformers are In-Context Molecular Designers
by: Liu, Gang, et al.
Published: (2025)
by: Liu, Gang, et al.
Published: (2025)
Hybrid Focal and Full-Range Attention Based Graph Transformers
by: Zhu, Minhong, et al.
Published: (2023)
by: Zhu, Minhong, et al.
Published: (2023)
Neighbourhood Transformer: Switchable Attention for Monophily-Aware Graph Learning
by: Luo, Yi, et al.
Published: (2026)
by: Luo, Yi, et al.
Published: (2026)
Twin Transformer using Gated Dynamic Learnable Attention mechanism for Fault Detection and Diagnosis in the Tennessee Eastman Process
by: Labbaf-Khaniki, Mohammad Ali, et al.
Published: (2024)
by: Labbaf-Khaniki, Mohammad Ali, et al.
Published: (2024)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
by: De Schouwer, Jonas, et al.
Published: (2026)
by: De Schouwer, Jonas, et al.
Published: (2026)
Continuous-Time Attention: PDE-Guided Mechanisms for Long-Sequence Transformers
by: Zhang, Yukun, et al.
Published: (2025)
by: Zhang, Yukun, et al.
Published: (2025)
Physics-Guided Transformer (PGT): Physics-Aware Attention Mechanism for PINNs
by: Zeraatkar, Ehsan, et al.
Published: (2026)
by: Zeraatkar, Ehsan, et al.
Published: (2026)
Enhancing Link Prediction with Fuzzy Graph Attention Networks and Dynamic Negative Sampling
by: Xing, Jinming, et al.
Published: (2024)
by: Xing, Jinming, et al.
Published: (2024)
RingFormer: A Ring-Enhanced Graph Transformer for Organic Solar Cell Property Prediction
by: Ding, Zhihao, et al.
Published: (2024)
by: Ding, Zhihao, et al.
Published: (2024)
Graph Attention-based Adaptive Transfer Learning for Link Prediction
by: Lu, Huashen, et al.
Published: (2025)
by: Lu, Huashen, et al.
Published: (2025)
Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
by: Dimitrov, Leon
Published: (2025)
by: Dimitrov, Leon
Published: (2025)
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
by: Wu, Gang, et al.
Published: (2025)
by: Wu, Gang, et al.
Published: (2025)
Spacetime $E(n)$-Transformer: Equivariant Attention for Spatio-temporal Graphs
by: Charles, Sergio G.
Published: (2024)
by: Charles, Sergio G.
Published: (2024)
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
by: Zhang, Jinhao, et al.
Published: (2026)
by: Zhang, Jinhao, et al.
Published: (2026)
Graph Diffusion Network for Drug-Gene Prediction
by: Wu, Jiayang, et al.
Published: (2025)
by: Wu, Jiayang, et al.
Published: (2025)
Discrete Diffusion Schrödinger Bridge Matching for Graph Transformation
by: Kim, Jun Hyeong, et al.
Published: (2024)
by: Kim, Jun Hyeong, et al.
Published: (2024)
VSFormer: Value and Shape-Aware Transformer with Prior-Enhanced Self-Attention for Multivariate Time Series Classification
by: Xi, Wenjie, et al.
Published: (2024)
by: Xi, Wenjie, et al.
Published: (2024)
Linear Transformers as VAR Models: Aligning Autoregressive Attention Mechanisms with Autoregressive Forecasting
by: Lu, Jiecheng, et al.
Published: (2025)
by: Lu, Jiecheng, et al.
Published: (2025)
Triple Attention Transformer Architecture for Time-Dependent Concrete Creep Prediction
by: Dokduea, Warayut, et al.
Published: (2025)
by: Dokduea, Warayut, et al.
Published: (2025)
TGFormer: Towards Temporal Graph Transformer with Auto-Correlation Mechanism
by: Chen, Hongjiang, et al.
Published: (2026)
by: Chen, Hongjiang, et al.
Published: (2026)
GraphCare: Enhancing Healthcare Predictions with Personalized Knowledge Graphs
by: Jiang, Pengcheng, et al.
Published: (2023)
by: Jiang, Pengcheng, et al.
Published: (2023)
Enhancing Linear Attention with Residual Learning
by: Lai, Xunhao, et al.
Published: (2025)
by: Lai, Xunhao, et al.
Published: (2025)
Cellular Traffic Prediction via Deep State Space Models with Attention Mechanism
by: Ma, Hui, et al.
Published: (2025)
by: Ma, Hui, et al.
Published: (2025)
Similar Items
-
Forgetting Transformer: Softmax Attention with a Forget Gate
by: Lin, Zhixuan, et al.
Published: (2025) -
A Simple yet Effective DDG Predictor is An Unsupervised Antibody Optimizer and Explainer
by: Wu, Lirong, et al.
Published: (2025) -
Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
by: Gupta, Abhijit
Published: (2026) -
Crossfusor: A Cross-Attention Transformer Enhanced Conditional Diffusion Model for Car-Following Trajectory Prediction
by: You, Junwei, et al.
Published: (2024) -
Value-State Gated Attention for Mitigating Extreme-Token Phenomena in Transformers
by: Bu, Rui, et al.
Published: (2025)