Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autor principal: Gupta, Abhijit
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918343521337344
author Gupta, Abhijit
author_facet Gupta, Abhijit
contents Molecular property prediction is crucial for drug discovery when labeled data are scarce. This work presents CardinalGraphFormer, a graph transformer augmented with a query-conditioned cardinality-preserving attention (CPA) channel that retains dynamic support-size signals complementary to static centrality embeddings. The approach combines structured sparse attention with Graphormer-inspired biases (shortest-path distance, centrality, direct-bond features) and unified dual-objective self-supervised pretraining (masked reconstruction and contrastive alignment of augmented views). Evaluation on 11 public benchmarks spanning MoleculeNet, OGB, and TDC ADMET demonstrates consistent improvements over protocol-matched baselines under matched pretraining, optimization, and hyperparameter tuning. Rigorous ablations confirm CPA's contributions and rule out simple size shortcuts. Code and reproducibility artifacts are provided.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02201
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
Gupta, Abhijit
Machine Learning
Artificial Intelligence
Molecular property prediction is crucial for drug discovery when labeled data are scarce. This work presents CardinalGraphFormer, a graph transformer augmented with a query-conditioned cardinality-preserving attention (CPA) channel that retains dynamic support-size signals complementary to static centrality embeddings. The approach combines structured sparse attention with Graphormer-inspired biases (shortest-path distance, centrality, direct-bond features) and unified dual-objective self-supervised pretraining (masked reconstruction and contrastive alignment of augmented views). Evaluation on 11 public benchmarks spanning MoleculeNet, OGB, and TDC ADMET demonstrates consistent improvements over protocol-matched baselines under matched pretraining, optimization, and hyperparameter tuning. Rigorous ablations confirm CPA's contributions and rule out simple size shortcuts. Code and reproducibility artifacts are provided.
title Cardinality-Preserving Attention Channels for Graph Transformers in Molecular Property Prediction
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2602.02201