GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pang, Yunhe, Chen, Bo, Zhang, Fanjin, Rao, Yanghui, Kharlamov, Evgeny, Tang, Jie
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866911094988079104
author Pang, Yunhe
Chen, Bo
Zhang, Fanjin
Rao, Yanghui
Kharlamov, Evgeny
Tang, Jie
author_facet Pang, Yunhe
Chen, Bo
Zhang, Fanjin
Rao, Yanghui
Kharlamov, Evgeny
Tang, Jie
contents Anomaly detection on text-rich graphs is widely prevalent in real life, such as detecting incorrectly assigned academic papers to authors and detecting bots in social networks. The remarkable capabilities of large language models (LLMs) pave a new revenue by utilizing rich-text information for effective anomaly detection. However, simply introducing rich texts into LLMs can obscure essential detection cues and introduce high fine-tuning costs. Moreover, LLMs often overlook the intrinsic structural bias of graphs which is vital for distinguishing normal from abnormal node patterns. To this end, this paper introduces GuARD, a text-rich and graph-informed language model that combines key structural features from graph-based methods with fine-grained semantic attributes extracted via small language models for effective anomaly detection on text-rich graphs. GuARD is optimized with the progressive multi-modal multi-turn instruction tuning framework in the task-guided instruction tuning regime tailed to incorporate both rich-text and structural modalities. Extensive experiments on four datasets reveal that GuARD outperforms graph-based and LLM-based anomaly detection methods, while offering up to 5$\times$ times speedup in training and 5$\times$ times speedup in inference over vanilla long-context LLMs on the large-scale WhoIsWho dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03930
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model
Pang, Yunhe
Chen, Bo
Zhang, Fanjin
Rao, Yanghui
Kharlamov, Evgeny
Tang, Jie
Computation and Language
Artificial Intelligence
Anomaly detection on text-rich graphs is widely prevalent in real life, such as detecting incorrectly assigned academic papers to authors and detecting bots in social networks. The remarkable capabilities of large language models (LLMs) pave a new revenue by utilizing rich-text information for effective anomaly detection. However, simply introducing rich texts into LLMs can obscure essential detection cues and introduce high fine-tuning costs. Moreover, LLMs often overlook the intrinsic structural bias of graphs which is vital for distinguishing normal from abnormal node patterns. To this end, this paper introduces GuARD, a text-rich and graph-informed language model that combines key structural features from graph-based methods with fine-grained semantic attributes extracted via small language models for effective anomaly detection on text-rich graphs. GuARD is optimized with the progressive multi-modal multi-turn instruction tuning framework in the task-guided instruction tuning regime tailed to incorporate both rich-text and structural modalities. Extensive experiments on four datasets reveal that GuARD outperforms graph-based and LLM-based anomaly detection methods, while offering up to 5$\times$ times speedup in training and 5$\times$ times speedup in inference over vanilla long-context LLMs on the large-scale WhoIsWho dataset.
title GuARD: Effective Anomaly Detection through a Text-Rich and Graph-Informed Language Model
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2412.03930