SigGate-GT: Taming Over-Smoothing in Graph Transformers via Sigmoid-Gated Attention
Fuente:
arXiv
Saved in:
| Main Authors: | Guo, Dongxin, Wu, Jikun, Yiu, Siu Ming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
Optimizing Inference in Transformer-Based Models: A Multi-Method Benchmark
by: Ho, Siu Hang, et al.
Published: (2025)
by: Ho, Siu Hang, et al.
Published: (2025)
Understanding and Tackling Over-Dilution in Graph Neural Networks
by: Lee, Junhyun, et al.
Published: (2025)
by: Lee, Junhyun, et al.
Published: (2025)
EvoPref: Multi-Objective Evolutionary Optimization Discovers Diverse LLM Alignments Beyond Gradient Descent
by: Guo, Dongxin, et al.
Published: (2026)
by: Guo, Dongxin, et al.
Published: (2026)
SCNode: Spatial and Contextual Coordinates for Graph Representation Learning
by: Uddin, Md Joshem, et al.
Published: (2024)
by: Uddin, Md Joshem, et al.
Published: (2024)
Persistent Topological Structures and Cohomological Flows as a Mathematical Framework for Brain-Inspired Representation Learning
by: Girish, Preksha, et al.
Published: (2025)
by: Girish, Preksha, et al.
Published: (2025)
Explicit Dropout: Deterministic Regularization for Transformer Architectures
by: Agrawal, Vidhi, et al.
Published: (2026)
by: Agrawal, Vidhi, et al.
Published: (2026)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
Expressivity of Representation Learning on Continuous-Time Dynamic Graphs: An Information-Flow Centric Review
by: Ennadir, Sofiane, et al.
Published: (2024)
by: Ennadir, Sofiane, et al.
Published: (2024)
TGraphX: Tensor-Aware Graph Neural Network for Multi-Dimensional Feature Learning
by: Sajjadi, Arash, et al.
Published: (2025)
by: Sajjadi, Arash, et al.
Published: (2025)
Massive Redundancy in Gradient Transport Enables Sparse Online Learning
by: Merin, Aur Shalev
Published: (2026)
by: Merin, Aur Shalev
Published: (2026)
Is Cambodia the World's Largest Cashew Producer?
by: Chaya, Veasna, et al.
Published: (2024)
by: Chaya, Veasna, et al.
Published: (2024)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
EchoLSTM: A Self-Reflective Recurrent Network for Stabilizing Long-Range Memory
by: K, Prasanth K, et al.
Published: (2025)
by: K, Prasanth K, et al.
Published: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
by: Qesaraku, Bjorna, et al.
Published: (2025)
by: Qesaraku, Bjorna, et al.
Published: (2025)
Complex-Valued Phase-Coherent Transformer
by: Hioki, Leona
Published: (2026)
by: Hioki, Leona
Published: (2026)
torchsom: The Reference PyTorch Library for Self-Organizing Maps
by: Berthier, Louis, et al.
Published: (2025)
by: Berthier, Louis, et al.
Published: (2025)
The Inhibitor: ReLU and Addition-Based Attention for Efficient Transformers under Fully Homomorphic Encryption on the Torus
by: Brännvall, Rickard, et al.
Published: (2023)
by: Brännvall, Rickard, et al.
Published: (2023)
Scaling Higher-Order Graph Learning with Maximal Clique Complexes
by: Vialle, Antoine, et al.
Published: (2026)
by: Vialle, Antoine, et al.
Published: (2026)
GLL: A Differentiable Graph Learning Layer for Neural Networks
by: Brown, Jason, et al.
Published: (2024)
by: Brown, Jason, et al.
Published: (2024)
Predicting Drug-Drug Interactions Using Heterogeneous Graph Neural Networks: HGNN-DDI
by: Liu, Hongbo, et al.
Published: (2025)
by: Liu, Hongbo, et al.
Published: (2025)
Tricks and Plug-ins for Gradient Boosting with Transformers
by: Fang, Biyi, et al.
Published: (2025)
by: Fang, Biyi, et al.
Published: (2025)
Just In Time Transformers
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
by: Benali, Ahmed Ala Eddine, et al.
Published: (2024)
Prediction-space knowledge markets for communication-efficient federated learning on multimedia tasks
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
Optimized Gradient Clipping for Noisy Label Learning
by: Ye, Xichen, et al.
Published: (2024)
by: Ye, Xichen, et al.
Published: (2024)
Contract-Driven QoE Auditing for Speech and Singing Services: From MOS Regression to Service Graphs
by: Du, Wenzhang
Published: (2025)
by: Du, Wenzhang
Published: (2025)
Leaf Spectral Reflectance Prediction Using Multi-Head Attention Neural Networks
by: Farajpoor, Parastoo, et al.
Published: (2026)
by: Farajpoor, Parastoo, et al.
Published: (2026)
The Normalized Difference Layer: A Differentiable Spectral Index Formulation for Deep Learning
by: Lotfi, Ali, et al.
Published: (2026)
by: Lotfi, Ali, et al.
Published: (2026)
H-Model: Dynamic Neural Architectures for Adaptive Processing
by: Hospodarchuk, Dmytro
Published: (2025)
by: Hospodarchuk, Dmytro
Published: (2025)
multivariateGPT: a decoder-only transformer for multivariate categorical and numeric data
by: Loza, Andrew J., et al.
Published: (2025)
by: Loza, Andrew J., et al.
Published: (2025)
VDW-GNNs: Vector diffusion wavelets for geometric graph neural networks
by: Johnson, David R., et al.
Published: (2025)
by: Johnson, David R., et al.
Published: (2025)
DecompKAN: Decomposed Patch-KAN for Long-Term Time Series Forecasting
by: Mysore, Naveen
Published: (2026)
by: Mysore, Naveen
Published: (2026)
Scaling Laws in the Tiny Regime: How Small Models Change Their Mistakes
by: Alnemari, Mohammed, et al.
Published: (2026)
by: Alnemari, Mohammed, et al.
Published: (2026)
Similar Items
-
Closing the Theory-Practice Gap in Spiking Transformers via Effective Dimension
by: Guo, Dongxin, et al.
Published: (2026) -
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
by: Guo, Dongxin, et al.
Published: (2026) -
When Do Early-Exit Networks Generalize? A PAC-Bayesian Theory of Adaptive Depth
by: Guo, Dongxin, et al.
Published: (2026) -
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
by: Guo, Dongxin, et al.
Published: (2026) -
Geometric Metrics for MoE Specialization: From Fisher Information to Early Failure Detection
by: Guo, Dongxin, et al.
Published: (2026)