SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Guo, Dongxin, Wu, Jikun, Yiu, Siu Ming
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908978143821824
author Guo, Dongxin
Wu, Jikun
Yiu, Siu Ming
author_facet Guo, Dongxin
Wu, Jikun
Yiu, Siu Ming
contents Graph transformers achieve strong results on molecular and long-range reasoning tasks, yet remain hampered by over-smoothing (the progressive collapse of node representations with depth) and attention entropy degeneration. We observe that these pathologies share a root cause with attention sinks in large language models: softmax attention's sum-to-one constraint forces every node to attend somewhere, even when no informative signal exists. Motivated by recent findings that element-wise sigmoid gating eliminates attention sinks in large language models, we propose SigGate-GT, a graph transformer that applies learned, per-head sigmoid gates to the attention output within the GraphGPS framework. Each gate can suppress activations toward zero, enabling heads to selectively silence uninformative connections. On five standard benchmarks, SigGate-GT matches the prior best on ZINC (0.059 MAE) and sets new state-of-the-art on ogbg-molhiv (82.47% ROC-AUC), with statistically significant gains over GraphGPS across all five datasets ($p < 0.05$). Ablations show that gating reduces over-smoothing by 30% (mean relative MAD gain across 4-16 layers), increases attention entropy, and stabilizes training across a $10\times$ learning rate range, with about 1% parameter overhead on OGB.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17324
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
Guo, Dongxin
Wu, Jikun
Yiu, Siu Ming
Machine Learning
Artificial Intelligence
68T07, 68R10
I.2.6; I.5.1; G.2.2
Graph transformers achieve strong results on molecular and long-range reasoning tasks, yet remain hampered by over-smoothing (the progressive collapse of node representations with depth) and attention entropy degeneration. We observe that these pathologies share a root cause with attention sinks in large language models: softmax attention's sum-to-one constraint forces every node to attend somewhere, even when no informative signal exists. Motivated by recent findings that element-wise sigmoid gating eliminates attention sinks in large language models, we propose SigGate-GT, a graph transformer that applies learned, per-head sigmoid gates to the attention output within the GraphGPS framework. Each gate can suppress activations toward zero, enabling heads to selectively silence uninformative connections. On five standard benchmarks, SigGate-GT matches the prior best on ZINC (0.059 MAE) and sets new state-of-the-art on ogbg-molhiv (82.47% ROC-AUC), with statistically significant gains over GraphGPS across all five datasets ($p < 0.05$). Ablations show that gating reduces over-smoothing by 30% (mean relative MAD gain across 4-16 layers), increases attention entropy, and stabilizes training across a $10\times$ learning rate range, with about 1% parameter overhead on OGB.
title SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
topic Machine Learning
Artificial Intelligence
68T07, 68R10
I.2.6; I.5.1; G.2.2
url https://arxiv.org/abs/2604.17324