The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Horych, Tomas, Mandl, Christoph, Ruas, Terry, Greiner-Petter, Andre, Gipp, Bela, Aizawa, Akiko, Spinde, Timo
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917901353615360
author Horych, Tomas
Mandl, Christoph
Ruas, Terry
Greiner-Petter, Andre
Gipp, Bela
Aizawa, Akiko
Spinde, Timo
author_facet Horych, Tomas
Mandl, Christoph
Ruas, Terry
Greiner-Petter, Andre
Gipp, Bela
Aizawa, Akiko
Spinde, Timo
contents High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality. LLMs have shown promising results in annotating downstream tasks like hate speech detection and political framing. Building on the success in these areas, this study investigates whether LLMs are viable for annotating the complex task of media bias detection and whether a downstream media bias classifier can be trained on such data. We create annolexical, the first large-scale dataset for media bias classification with over 48000 synthetically annotated examples. Our classifier, fine-tuned on this dataset, surpasses all of the annotator LLMs by 5-9 percent in Matthews Correlation Coefficient (MCC) and performs close to or outperforms the model trained on human-labeled data when evaluated on two media bias benchmark datasets (BABE and BASIL). This study demonstrates how our approach significantly reduces the cost of dataset creation in the media bias domain and, by extension, the development of classifiers, while our subsequent behavioral stress-testing reveals some of its current limitations and trade-offs.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11081
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
Horych, Tomas
Mandl, Christoph
Ruas, Terry
Greiner-Petter, Andre
Gipp, Bela
Aizawa, Akiko
Spinde, Timo
Computation and Language
High annotation costs from hiring or crowdsourcing complicate the creation of large, high-quality datasets needed for training reliable text classifiers. Recent research suggests using Large Language Models (LLMs) to automate the annotation process, reducing these costs while maintaining data quality. LLMs have shown promising results in annotating downstream tasks like hate speech detection and political framing. Building on the success in these areas, this study investigates whether LLMs are viable for annotating the complex task of media bias detection and whether a downstream media bias classifier can be trained on such data. We create annolexical, the first large-scale dataset for media bias classification with over 48000 synthetically annotated examples. Our classifier, fine-tuned on this dataset, surpasses all of the annotator LLMs by 5-9 percent in Matthews Correlation Coefficient (MCC) and performs close to or outperforms the model trained on human-labeled data when evaluated on two media bias benchmark datasets (BABE and BASIL). This study demonstrates how our approach significantly reduces the cost of dataset creation in the media bias domain and, by extension, the development of classifiers, while our subsequent behavioral stress-testing reveals some of its current limitations and trade-offs.
title The Promises and Pitfalls of LLM Annotations in Dataset Labeling: a Case Study on Media Bias Detection
topic Computation and Language
url https://arxiv.org/abs/2411.11081