What Improves the Generalization of Graph Transformers? A Theoretical Dive into the Self-attention and Positional Encoding
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Hongkang, Wang, Meng, Ma, Tengfei, Liu, Sijia, Zhang, Zaixi, Chen, Pin-Yu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Improving Transformers using Faithful Positional Encoding
by: Idé, Tsuyoshi, et al.
Published: (2024)
by: Idé, Tsuyoshi, et al.
Published: (2024)
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
How does promoting the minority fraction affect generalization? A theoretical study of the one-hidden-layer neural network on group imbalance
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
How Do Nonlinear Transformers Learn and Generalize in In-Context Learning?
by: Li, Hongkang, et al.
Published: (2024)
by: Li, Hongkang, et al.
Published: (2024)
Theoretical Learning Performance of Graph Neural Networks: The Impact of Jumping Connections and Layer-wise Sparsification
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
Enhancing Graph Transformers with Hierarchical Distance Structural Encoding
by: Luo, Yuankai, et al.
Published: (2023)
by: Luo, Yuankai, et al.
Published: (2023)
Graph Transformer with Disease Subgraph Positional Encoding for Improved Comorbidity Prediction
by: Qin, Xihan, et al.
Published: (2025)
by: Qin, Xihan, et al.
Published: (2025)
What Are Good Positional Encodings for Directed Graphs?
by: Huang, Yinan, et al.
Published: (2024)
by: Huang, Yinan, et al.
Published: (2024)
FedGT: Federated Node Classification with Scalable Graph Transformer
by: Zhang, Zaixi, et al.
Published: (2024)
by: Zhang, Zaixi, et al.
Published: (2024)
Visual prompting reimagined: The power of the Activation Prompts
by: Zhang, Yihua, et al.
Published: (2026)
by: Zhang, Yihua, et al.
Published: (2026)
A Theoretical Analysis of Mamba's Training Dynamics: Filtering Relevant Features for Generalization in State Space Models
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
by: Shandirasegaran, Mugunthan, et al.
Published: (2026)
Graph Diffusion Transformers for Multi-Conditional Molecular Generation
by: Liu, Gang, et al.
Published: (2024)
by: Liu, Gang, et al.
Published: (2024)
Dynamic Graph Transformer with Correlated Spatial-Temporal Positional Encoding
by: Wang, Zhe, et al.
Published: (2024)
by: Wang, Zhe, et al.
Published: (2024)
Comparing Graph Transformers via Positional Encodings
by: Black, Mitchell, et al.
Published: (2024)
by: Black, Mitchell, et al.
Published: (2024)
Graph Transformers without Positional Encodings
by: Garg, Ayush
Published: (2024)
by: Garg, Ayush
Published: (2024)
Size Transferability of Graph Transformers with Convolutional Positional Encodings
by: Porras-Valenzuela, Javier, et al.
Published: (2026)
by: Porras-Valenzuela, Javier, et al.
Published: (2026)
Benchmarking Positional Encodings for GNNs and Graph Transformers
by: Grötschla, Florian, et al.
Published: (2024)
by: Grötschla, Florian, et al.
Published: (2024)
On the Stability of Expressive Positional Encodings for Graphs
by: Huang, Yinan, et al.
Published: (2023)
by: Huang, Yinan, et al.
Published: (2023)
Towards Few-shot Self-explaining Graph Neural Networks
by: Peng, Jingyu, et al.
Published: (2024)
by: Peng, Jingyu, et al.
Published: (2024)
Tab-PET: Graph-Based Positional Encodings for Tabular Transformers
by: Leng, Yunze, et al.
Published: (2025)
by: Leng, Yunze, et al.
Published: (2025)
DAM-GT: Dual Positional Encoding-Based Attention Masking Graph Transformer for Node Classification
by: Li, Chenyang, et al.
Published: (2025)
by: Li, Chenyang, et al.
Published: (2025)
Vocabulary In-Context Learning in Transformers: Benefits of Positional Encoding
by: Ma, Qian, et al.
Published: (2025)
by: Ma, Qian, et al.
Published: (2025)
Theoretical Analysis of Positional Encodings in Transformer Models: Impact on Expressiveness and Generalization
by: Li, Yin
Published: (2025)
by: Li, Yin
Published: (2025)
Learning Laplacian Positional Encodings for Heterophilous Graphs
by: Ito, Michael, et al.
Published: (2025)
by: Ito, Michael, et al.
Published: (2025)
Graph Self-Supervised Learning with Learnable Structural and Positional Encodings
by: Wijesinghe, Asiri, et al.
Published: (2025)
by: Wijesinghe, Asiri, et al.
Published: (2025)
Transformers Learn the Optimal DDPM Denoiser for Multi-Token GMMs
by: Li, Hongkang, et al.
Published: (2026)
by: Li, Hongkang, et al.
Published: (2026)
On the Geometry of Positional Encodings in Transformers
by: Cirrincione, Giansalvo
Published: (2026)
by: Cirrincione, Giansalvo
Published: (2026)
Rethinking Time Encoding via Learnable Transformation Functions
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Position Encoding with Random Float Sampling Enhances Length Generalization of Transformers
by: Shimizu, Atsushi, et al.
Published: (2026)
by: Shimizu, Atsushi, et al.
Published: (2026)
Rotary Position Encodings for Graphs
by: Reid, Isaac, et al.
Published: (2025)
by: Reid, Isaac, et al.
Published: (2025)
Improving Position Encoding of Transformers for Multivariate Time Series Classification
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
by: Foumani, Navid Mohammadi, et al.
Published: (2023)
Learning a Fourier Transform for Linear Relative Positional Encodings in Transformers
by: Choromanski, Krzysztof Marcin, et al.
Published: (2023)
by: Choromanski, Krzysztof Marcin, et al.
Published: (2023)
PSBD: Prediction Shift Uncertainty Unlocks Backdoor Detection
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Graph Diffusion Transformers are In-Context Molecular Designers
by: Liu, Gang, et al.
Published: (2025)
by: Liu, Gang, et al.
Published: (2025)
Towards Graph Foundation Models: A Study on the Generalization of Positional and Structural Encodings
by: Franks, Billy Joe, et al.
Published: (2024)
by: Franks, Billy Joe, et al.
Published: (2024)
Agentic-VLA: Efficient Online Adaptation for Vision-Language-Action Models
by: Jin, Ruofan, et al.
Published: (2026)
by: Jin, Ruofan, et al.
Published: (2026)
SaGIF: Improving Individual Fairness in Graph Neural Networks via Similarity Encoding
by: Zhu, Yuchang, et al.
Published: (2025)
by: Zhu, Yuchang, et al.
Published: (2025)
Similar Items
-
Learning on Transformers is Provable Low-Rank and Sparse: A One-layer Analysis
by: Li, Hongkang, et al.
Published: (2024) -
When is Task Vector Provably Effective for Model Editing? A Generalization Analysis of Nonlinear Transformers
by: Li, Hongkang, et al.
Published: (2025) -
Improving Transformers using Faithful Positional Encoding
by: Idé, Tsuyoshi, et al.
Published: (2024) -
Training Nonlinear Transformers for Chain-of-Thought Inference: A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2024) -
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)