SQBC: Active Learning using LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Wagner, Stefan Sylvius, Behrendt, Maike, Ziegele, Marc, Harmeling, Stefan
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913311120949248
author Wagner, Stefan Sylvius
Behrendt, Maike
Ziegele, Marc
Harmeling, Stefan
author_facet Wagner, Stefan Sylvius
Behrendt, Maike
Ziegele, Marc
Harmeling, Stefan
contents Stance detection is an important task for many applications that analyse or support online political discussions. Common approaches include fine-tuning transformer based models. However, these models require a large amount of labelled data, which might not be available. In this work, we present two different ways to leverage LLM-generated synthetic data to train and improve stance detection agents for online political discussions: first, we show that augmenting a small fine-tuning dataset with synthetic data can improve the performance of the stance detection model. Second, we propose a new active learning method called SQBC based on the "Query-by-Comittee" approach. The key idea is to use LLM-generated synthetic data as an oracle to identify the most informative unlabelled samples, that are selected for manual labelling. Comprehensive experiments show that both ideas can improve the stance detection performance. Curiously, we observed that fine-tuning on actively selected samples can exceed the performance of using the full dataset.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08078
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle SQBC: Active Learning using LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions
Wagner, Stefan Sylvius
Behrendt, Maike
Ziegele, Marc
Harmeling, Stefan
Computation and Language
Artificial Intelligence
Machine Learning
Stance detection is an important task for many applications that analyse or support online political discussions. Common approaches include fine-tuning transformer based models. However, these models require a large amount of labelled data, which might not be available. In this work, we present two different ways to leverage LLM-generated synthetic data to train and improve stance detection agents for online political discussions: first, we show that augmenting a small fine-tuning dataset with synthetic data can improve the performance of the stance detection model. Second, we propose a new active learning method called SQBC based on the "Query-by-Comittee" approach. The key idea is to use LLM-generated synthetic data as an oracle to identify the most informative unlabelled samples, that are selected for manual labelling. Comprehensive experiments show that both ideas can improve the stance detection performance. Curiously, we observed that fine-tuning on actively selected samples can exceed the performance of using the full dataset.
title SQBC: Active Learning using LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.08078