ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Brannon, William, Kang, Wonjune, Fulay, Suyash, Jiang, Hang, Roy, Brandon, Roy, Deb, Kabbara, Jad
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909248843153408
author Brannon, William
Kang, Wonjune
Fulay, Suyash
Jiang, Hang
Roy, Brandon
Roy, Deb
Kabbara, Jad
author_facet Brannon, William
Kang, Wonjune
Fulay, Suyash
Jiang, Hang
Roy, Brandon
Roy, Deb
Kabbara, Jad
contents Learning on text-attributed graphs (TAGs), in which nodes are associated with one or more texts, has been the subject of much recent work. However, most approaches tend to make strong assumptions about the downstream task of interest, are reliant on hand-labeled data, or fail to equally balance the importance of both text and graph representations. In this work, we propose Contrastive Graph-Text pretraining (ConGraT), a general, self-supervised approach for jointly learning separate representations of texts and nodes in a TAG. Our method trains a language model (LM) and a graph neural network (GNN) to align their representations in a common latent space using a batch-wise contrastive learning objective inspired by CLIP. We further propose an extension to the CLIP objective that leverages graph structure to incorporate information about inter-node similarity. Extensive experiments demonstrate that ConGraT outperforms baselines on various downstream tasks, including node and text category classification, link prediction, and language modeling. Finally, we present an application of our method to community detection in social graphs, which enables finding more textually grounded communities, rather than purely graph-based ones. Code and certain datasets are available at https://github.com/wwbrannon/congrat.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14321
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings
Brannon, William
Kang, Wonjune
Fulay, Suyash
Jiang, Hang
Roy, Brandon
Roy, Deb
Kabbara, Jad
Computation and Language
Learning on text-attributed graphs (TAGs), in which nodes are associated with one or more texts, has been the subject of much recent work. However, most approaches tend to make strong assumptions about the downstream task of interest, are reliant on hand-labeled data, or fail to equally balance the importance of both text and graph representations. In this work, we propose Contrastive Graph-Text pretraining (ConGraT), a general, self-supervised approach for jointly learning separate representations of texts and nodes in a TAG. Our method trains a language model (LM) and a graph neural network (GNN) to align their representations in a common latent space using a batch-wise contrastive learning objective inspired by CLIP. We further propose an extension to the CLIP objective that leverages graph structure to incorporate information about inter-node similarity. Extensive experiments demonstrate that ConGraT outperforms baselines on various downstream tasks, including node and text category classification, link prediction, and language modeling. Finally, we present an application of our method to community detection in social graphs, which enables finding more textually grounded communities, rather than purely graph-based ones. Code and certain datasets are available at https://github.com/wwbrannon/congrat.
title ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings
topic Computation and Language
url https://arxiv.org/abs/2305.14321