A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Verhoeven, Ivo, Mishra, Pushkar, Beloch, Rahel, Yannakoudakis, Helen, Shutova, Ekaterina
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909327466430464
author Verhoeven, Ivo
Mishra, Pushkar
Beloch, Rahel
Yannakoudakis, Helen
Shutova, Ekaterina
author_facet Verhoeven, Ivo
Mishra, Pushkar
Beloch, Rahel
Yannakoudakis, Helen
Shutova, Ekaterina
contents Community models for malicious content detection, which take into account the context from a social graph alongside the content itself, have shown remarkable performance on benchmark datasets. Yet, misinformation and hate speech continue to propagate on social media networks. This mismatch can be partially attributed to the limitations of current evaluation setups that neglect the rapid evolution of online content and the underlying social graph. In this paper, we propose a novel evaluation setup for model generalisation based on our few-shot subgraph sampling approach. This setup tests for generalisation through few labelled examples in local explorations of a larger graph, emulating more realistic application settings. We show this to be a challenging inductive setup, wherein strong performance on the training graph is not indicative of performance on unseen tasks, domains, or graph structures. Lastly, we show that graph meta-learners trained with our proposed few-shot subgraph sampling outperform standard community models in the inductive setup. We make our code publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2404_01822
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
Verhoeven, Ivo
Mishra, Pushkar
Beloch, Rahel
Yannakoudakis, Helen
Shutova, Ekaterina
Machine Learning
Computation and Language
Social and Information Networks
Community models for malicious content detection, which take into account the context from a social graph alongside the content itself, have shown remarkable performance on benchmark datasets. Yet, misinformation and hate speech continue to propagate on social media networks. This mismatch can be partially attributed to the limitations of current evaluation setups that neglect the rapid evolution of online content and the underlying social graph. In this paper, we propose a novel evaluation setup for model generalisation based on our few-shot subgraph sampling approach. This setup tests for generalisation through few labelled examples in local explorations of a larger graph, emulating more realistic application settings. We show this to be a challenging inductive setup, wherein strong performance on the training graph is not indicative of performance on unseen tasks, domains, or graph structures. Lastly, we show that graph meta-learners trained with our proposed few-shot subgraph sampling outperform standard community models in the inductive setup. We make our code publicly available.
title A (More) Realistic Evaluation Setup for Generalisation of Community Models on Malicious Content Detection
topic Machine Learning
Computation and Language
Social and Information Networks
url https://arxiv.org/abs/2404.01822