Scaling Graph Transformers: A Comparative Study of Sparse and Dense Attention
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Dimitrov, Leon |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Generalizing Scaling Laws for Dense and Sparse Large Language Models
par: Hossain, Md Arafat, et autres
Publié: (2025)
par: Hossain, Md Arafat, et autres
Publié: (2025)
A Comparative Study on Dynamic Graph Embedding based on Mamba and Transformers
par: Pandey, Ashish Parmanand, et autres
Publié: (2024)
par: Pandey, Ashish Parmanand, et autres
Publié: (2024)
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
par: Wu, Gang, et autres
Publié: (2025)
par: Wu, Gang, et autres
Publié: (2025)
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
par: Bouadi, Mohamed, et autres
Publié: (2025)
par: Bouadi, Mohamed, et autres
Publié: (2025)
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
par: El, Batu, et autres
Publié: (2025)
par: El, Batu, et autres
Publié: (2025)
Graph Convolutions Enrich the Self-Attention in Transformers!
par: Choi, Jeongwhan, et autres
Publié: (2023)
par: Choi, Jeongwhan, et autres
Publié: (2023)
FlashOmni: A Unified Sparse Attention Engine for Diffusion Transformers
par: Qiao, Liang, et autres
Publié: (2025)
par: Qiao, Liang, et autres
Publié: (2025)
vAttention: Verified Sparse Attention
par: Desai, Aditya, et autres
Publié: (2025)
par: Desai, Aditya, et autres
Publié: (2025)
Self-Tuning Sparse Attention: Multi-Fidelity Hyperparameter Optimization for Transformer Acceleration
par: Dev, Arundhathi, et autres
Publié: (2026)
par: Dev, Arundhathi, et autres
Publié: (2026)
Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
par: Panda, Ashwinee, et autres
Publié: (2025)
par: Panda, Ashwinee, et autres
Publié: (2025)
InfLLM-V2: Dense-Sparse Switchable Attention for Seamless Short-to-Long Adaptation
par: Zhao, Weilin, et autres
Publié: (2025)
par: Zhao, Weilin, et autres
Publié: (2025)
VecFormer: Towards Efficient and Generalizable Graph Transformer with Graph Token Attention
par: Zhou, Jingbo, et autres
Publié: (2026)
par: Zhou, Jingbo, et autres
Publié: (2026)
TQL: Scaling Q-Functions with Transformers by Preventing Attention Collapse
par: Dong, Perry, et autres
Publié: (2026)
par: Dong, Perry, et autres
Publié: (2026)
Nexusformer: Nonlinear Attention Expansion for Stable and Inheritable Transformer Scaling
par: Zhao, Weijie, et autres
Publié: (2026)
par: Zhao, Weijie, et autres
Publié: (2026)
Scaling Laws Across Model Architectures: A Comparative Analysis of Dense and MoE Models in Large Language Models
par: Wang, Siqi, et autres
Publié: (2024)
par: Wang, Siqi, et autres
Publié: (2024)
Graph Attention-Guided Search for Dense Multi-Agent Pathfinding
par: Jain, Rishabh, et autres
Publié: (2025)
par: Jain, Rishabh, et autres
Publié: (2025)
Hybrid Focal and Full-Range Attention Based Graph Transformers
par: Zhu, Minhong, et autres
Publié: (2023)
par: Zhu, Minhong, et autres
Publié: (2023)
Neighbourhood Transformer: Switchable Attention for Monophily-Aware Graph Learning
par: Luo, Yi, et autres
Publié: (2026)
par: Luo, Yi, et autres
Publié: (2026)
k-Maximum Inner Product Attention for Graph Transformers and the Expressive Power of GraphGPS
par: De Schouwer, Jonas, et autres
Publié: (2026)
par: De Schouwer, Jonas, et autres
Publié: (2026)
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
par: Gao, Yizhao, et autres
Publié: (2025)
par: Gao, Yizhao, et autres
Publié: (2025)
Spacetime $E(n)$-Transformer: Equivariant Attention for Spatio-temporal Graphs
par: Charles, Sergio G.
Publié: (2024)
par: Charles, Sergio G.
Publié: (2024)
Attention Dispersion in Dynamic Graph Transformers: Diagnosis and a Transferable Fix
par: Zhang, Jinhao, et autres
Publié: (2026)
par: Zhang, Jinhao, et autres
Publié: (2026)
Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
par: Gupta, Abhijit
Publié: (2026)
par: Gupta, Abhijit
Publié: (2026)
Improving Sparse Autoencoder with Dynamic Attention
par: Wang, Dongsheng, et autres
Publié: (2026)
par: Wang, Dongsheng, et autres
Publié: (2026)
GraphSparseNet: a Novel Method for Large Scale Traffic Flow Prediction
par: Kong, Weiyang, et autres
Publié: (2025)
par: Kong, Weiyang, et autres
Publié: (2025)
Beyond Dense States: Elevating Sparse Transcoders to Active Operators for Latent Reasoning
par: Wang, Yadong, et autres
Publié: (2026)
par: Wang, Yadong, et autres
Publié: (2026)
SparseBalance: Load-Balanced Long Context Training with Dynamic Sparse Attention
par: Xu, Hongtao, et autres
Publié: (2026)
par: Xu, Hongtao, et autres
Publié: (2026)
A Comparative Study of Specialized LLMs as Dense Retrievers
par: Zhang, Hengran, et autres
Publié: (2025)
par: Zhang, Hengran, et autres
Publié: (2025)
EquiformerV3: Scaling Efficient, Expressive, and General SE(3)-Equivariant Graph Attention Transformers
par: Liao, Yi-Lun, et autres
Publié: (2026)
par: Liao, Yi-Lun, et autres
Publié: (2026)
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
par: Yuan, Jingyang, et autres
Publié: (2025)
par: Yuan, Jingyang, et autres
Publié: (2025)
Longer Attention Span: Increasing Transformer Context Length with Sparse Graph Processing Techniques
par: Tomczak, Nathaniel, et autres
Publié: (2025)
par: Tomczak, Nathaniel, et autres
Publié: (2025)
Sparse Probabilistic Graph Circuits
par: Rektoris, Martin, et autres
Publié: (2025)
par: Rektoris, Martin, et autres
Publié: (2025)
How Sparse Attention Approximates Exact Attention? Your Attention is Naturally $n^C$-Sparse
par: Deng, Yichuan, et autres
Publié: (2024)
par: Deng, Yichuan, et autres
Publié: (2024)
DGTN: Graph-Enhanced Transformer with Diffusive Attention Gating Mechanism for Enzyme DDG Prediction
par: Lin, Abigail
Publié: (2025)
par: Lin, Abigail
Publié: (2025)
BiScale-GTR: Fragment-Aware Graph Transformers for Multi-Scale Molecular Representation Learning
par: Yang, Yi, et autres
Publié: (2026)
par: Yang, Yi, et autres
Publié: (2026)
CVTGAD: Simplified Transformer with Cross-View Attention for Unsupervised Graph-level Anomaly Detection
par: Li, Jindong, et autres
Publié: (2024)
par: Li, Jindong, et autres
Publié: (2024)
B-TGAT: A Bi-directional Temporal Graph Attention Transformer for Clustering Multivariate Spatiotemporal Data
par: Nji, Francis Ndikum, et autres
Publié: (2025)
par: Nji, Francis Ndikum, et autres
Publié: (2025)
Multi-Scale Adaptive Neighborhood Awareness Transformer For Graph Fraud Detection
par: Lv, Jiaqi, et autres
Publié: (2026)
par: Lv, Jiaqi, et autres
Publié: (2026)
Progressive Sparse Attention: Algorithm and System Co-design for Efficient Attention in LLM Serving
par: Zhou, Qihui, et autres
Publié: (2025)
par: Zhou, Qihui, et autres
Publié: (2025)
Attention in Constant Time: Vashista Sparse Attention for Long-Context Decoding with Exponential Guarantees
par: Nobaub, Vashista
Publié: (2026)
par: Nobaub, Vashista
Publié: (2026)
Documents similaires
-
Generalizing Scaling Laws for Dense and Sparse Large Language Models
par: Hossain, Md Arafat, et autres
Publié: (2025) -
A Comparative Study on Dynamic Graph Embedding based on Mamba and Transformers
par: Pandey, Ashish Parmanand, et autres
Publié: (2024) -
Exploring the Global-to-Local Attention Scheme in Graph Transformers: An Empirical Study
par: Wu, Gang, et autres
Publié: (2025) -
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
par: Bouadi, Mohamed, et autres
Publié: (2025) -
Towards Mechanistic Interpretability of Graph Transformers via Attention Graphs
par: El, Batu, et autres
Publié: (2025)