Masked Graph Transformer for Large-Scale Recommendation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866909192754823168 |
|---|---|
| author | Chen, Huiyuan Xu, Zhe Yeh, Chin-Chia Michael Lai, Vivian Zheng, Yan Xu, Minghua Tong, Hanghang |
| author_facet | Chen, Huiyuan Xu, Zhe Yeh, Chin-Chia Michael Lai, Vivian Zheng, Yan Xu, Minghua Tong, Hanghang |
| contents | Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the scalability of Graph Transformers, particularly for large-scale recommendation. Here we propose an efficient Masked Graph Transformer, named MGFormer, capable of capturing all-pair interactions among nodes with a linear complexity. To achieve this, we treat all user/item nodes as independent tokens, enhance them with positional embeddings, and feed them into a kernelized attention module. Additionally, we incorporate learnable relative degree information to appropriately reweigh the attentions. Experimental results show the superior performance of our MGFormer, even with a single attention layer. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2405_04028 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Masked Graph Transformer for Large-Scale Recommendation Chen, Huiyuan Xu, Zhe Yeh, Chin-Chia Michael Lai, Vivian Zheng, Yan Xu, Minghua Tong, Hanghang Information Retrieval Graph Transformers have garnered significant attention for learning graph-structured data, thanks to their superb ability to capture long-range dependencies among nodes. However, the quadratic space and time complexity hinders the scalability of Graph Transformers, particularly for large-scale recommendation. Here we propose an efficient Masked Graph Transformer, named MGFormer, capable of capturing all-pair interactions among nodes with a linear complexity. To achieve this, we treat all user/item nodes as independent tokens, enhance them with positional embeddings, and feed them into a kernelized attention module. Additionally, we incorporate learnable relative degree information to appropriately reweigh the attentions. Experimental results show the superior performance of our MGFormer, even with a single attention layer. |
| title | Masked Graph Transformer for Large-Scale Recommendation |
| topic | Information Retrieval |
| url | https://arxiv.org/abs/2405.04028 |