Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909524415217664 |
|---|---|
| author | Choi, Eun Cheol Balasubramanian, Ashwin Qi, Jinhu Ferrara, Emilio |
| author_facet | Choi, Eun Cheol Balasubramanian, Ashwin Qi, Jinhu Ferrara, Emilio |
| contents | Misinformation surrounding emerging outbreaks poses a serious societal threat, making robust countermeasures essential. One promising approach is stance detection (SD), which identifies whether social media posts support or oppose misleading claims. In this work, we finetune classifiers on COVID-19 misinformation SD datasets consisting of claims and corresponding tweets. Specifically, we test controllable misinformation generation (CMG) using large language models (LLMs) as a method for data augmentation. While CMG demonstrates the potential for expanding training datasets, our experiments reveal that performance gains over traditional augmentation methods are often minimal and inconsistent, primarily due to built-in safeguards within LLMs. We release our code and datasets to facilitate further research on misinformation detection and generation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_02328 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection Choi, Eun Cheol Balasubramanian, Ashwin Qi, Jinhu Ferrara, Emilio Computation and Language Computers and Society Human-Computer Interaction Social and Information Networks Misinformation surrounding emerging outbreaks poses a serious societal threat, making robust countermeasures essential. One promising approach is stance detection (SD), which identifies whether social media posts support or oppose misleading claims. In this work, we finetune classifiers on COVID-19 misinformation SD datasets consisting of claims and corresponding tweets. Specifically, we test controllable misinformation generation (CMG) using large language models (LLMs) as a method for data augmentation. While CMG demonstrates the potential for expanding training datasets, our experiments reveal that performance gains over traditional augmentation methods are often minimal and inconsistent, primarily due to built-in safeguards within LLMs. We release our code and datasets to facilitate further research on misinformation detection and generation. |
| title | Limited Effectiveness of LLM-based Data Augmentation for COVID-19 Misinformation Stance Detection |
| topic | Computation and Language Computers and Society Human-Computer Interaction Social and Information Networks |
| url | https://arxiv.org/abs/2503.02328 |