Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Choi, Eun Cheol, Balasubramanian, Ashwin, Qi, Jinhu, Ferrara, Emilio
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909524415217664
author Choi, Eun Cheol
Balasubramanian, Ashwin
Qi, Jinhu
Ferrara, Emilio
author_facet Choi, Eun Cheol
Balasubramanian, Ashwin
Qi, Jinhu
Ferrara, Emilio
contents Misinformation surrounding emerging outbreaks poses a serious societal threat, making robust countermeasures essential. One promising approach is stance detection (SD), which identifies whether social media posts support or oppose misleading claims. In this work, we finetune classifiers on COVID-19 misinformation SD datasets consisting of claims and corresponding tweets. Specifically, we test controllable misinformation generation (CMG) using large language models (LLMs) as a method for data augmentation. While CMG demonstrates the potential for expanding training datasets, our experiments reveal that performance gains over traditional augmentation methods are often minimal and inconsistent, primarily due to built-in safeguards within LLMs. We release our code and datasets to facilitate further research on misinformation detection and generation.
format Preprint
id arxiv_https___arxiv_org_abs_2503_02328
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection
Choi, Eun Cheol
Balasubramanian, Ashwin
Qi, Jinhu
Ferrara, Emilio
Computation and Language
Computers and Society
Human-Computer Interaction
Social and Information Networks
Misinformation surrounding emerging outbreaks poses a serious societal threat, making robust countermeasures essential. One promising approach is stance detection (SD), which identifies whether social media posts support or oppose misleading claims. In this work, we finetune classifiers on COVID-19 misinformation SD datasets consisting of claims and corresponding tweets. Specifically, we test controllable misinformation generation (CMG) using large language models (LLMs) as a method for data augmentation. While CMG demonstrates the potential for expanding training datasets, our experiments reveal that performance gains over traditional augmentation methods are often minimal and inconsistent, primarily due to built-in safeguards within LLMs. We release our code and datasets to facilitate further research on misinformation detection and generation.
title Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection
topic Computation and Language
Computers and Society
Human-Computer Interaction
Social and Information Networks
url https://arxiv.org/abs/2503.02328