Masked Graph Transformer for Large-Scale Recommendation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Huiyuan, Xu, Zhe, Yeh, Chin-Chia Michael, Lai, Vivian, Zheng, Yan, Xu, Minghua, Tong, Hanghang
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909192754823168
author Chen, Huiyuan
Xu, Zhe
Yeh, Chin-Chia Michael
Lai, Vivian
Zheng, Yan
Xu, Minghua
Tong, Hanghang
author_facet Chen, Huiyuan
Xu, Zhe
Yeh, Chin-Chia Michael
Lai, Vivian
Zheng, Yan
Xu, Minghua
Tong, Hanghang
contents Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the scalability of Graph Transformers, particularly for large-scale recommendation. Here we propose an efficient Masked Graph Transformer, named MGFormer, capable of capturing all-pair interactions among nodes with a linear complexity. To achieve this, we treat all user/item nodes as independent tokens, enhance them with positional embeddings, and feed them into a kernelized attention module. Additionally, we incorporate learnable relative degree information to appropriately reweigh the attentions. Experimental results show the superior performance of our MGFormer, even with a single attention layer.
format Preprint
id arxiv_https___arxiv_org_abs_2405_04028
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Masked Graph Transformer for Large-Scale Recommendation
Chen, Huiyuan
Xu, Zhe
Yeh, Chin-Chia Michael
Lai, Vivian
Zheng, Yan
Xu, Minghua
Tong, Hanghang
Information Retrieval
Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the scalability of Graph Transformers, particularly for large-scale recommendation. Here we propose an efficient Masked Graph Transformer, named MGFormer, capable of capturing all-pair interactions among nodes with a linear complexity. To achieve this, we treat all user/item nodes as independent tokens, enhance them with positional embeddings, and feed them into a kernelized attention module. Additionally, we incorporate learnable relative degree information to appropriately reweigh the attentions. Experimental results show the superior performance of our MGFormer, even with a single attention layer.
title Masked Graph Transformer for Large-Scale Recommendation
topic Information Retrieval
url https://arxiv.org/abs/2405.04028