The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Møller, Anders Giovanni, Dalsgaard, Jacob Aarup, Pera, Arianna, Aiello, Luca Maria
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911770512195584
author Møller, Anders Giovanni
Dalsgaard, Jacob Aarup
Pera, Arianna
Aiello, Luca Maria
author_facet Møller, Anders Giovanni
Dalsgaard, Jacob Aarup
Pera, Arianna
Aiello, Luca Maria
contents In the realm of Computational Social Science (CSS), practitioners often navigate complex, low-resource domains and face the costly and time-intensive challenges of acquiring and annotating data. We aim to establish a set of guidelines to address such challenges, comparing the use of human-labeled data with synthetically generated data from GPT-4 and Llama-2 in ten distinct CSS classification tasks of varying complexity. Additionally, we examine the impact of training data sizes on performance. Our findings reveal that models trained on human-labeled data consistently exhibit superior or comparable performance compared to their synthetically augmented counterparts. Nevertheless, synthetic augmentation proves beneficial, particularly in improving performance on rare classes within multi-class tasks. Furthermore, we leverage GPT-4 and Llama-2 for zero-shot classification and find that, while they generally display strong performance, they often fall short when compared to specialized classifiers trained on moderately sized training sets.
format Preprint
id arxiv_https___arxiv_org_abs_2304_13861
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
Møller, Anders Giovanni
Dalsgaard, Jacob Aarup
Pera, Arianna
Aiello, Luca Maria
Computation and Language
Computers and Society
Physics and Society
In the realm of Computational Social Science (CSS), practitioners often navigate complex, low-resource domains and face the costly and time-intensive challenges of acquiring and annotating data. We aim to establish a set of guidelines to address such challenges, comparing the use of human-labeled data with synthetically generated data from GPT-4 and Llama-2 in ten distinct CSS classification tasks of varying complexity. Additionally, we examine the impact of training data sizes on performance. Our findings reveal that models trained on human-labeled data consistently exhibit superior or comparable performance compared to their synthetically augmented counterparts. Nevertheless, synthetic augmentation proves beneficial, particularly in improving performance on rare classes within multi-class tasks. Furthermore, we leverage GPT-4 and Llama-2 for zero-shot classification and find that, while they generally display strong performance, they often fall short when compared to specialized classifiers trained on moderately sized training sets.
title The Parrot Dilemma: Human-Labeled vs. LLM-augmented Data in Classification Tasks
topic Computation and Language
Computers and Society
Physics and Society
url https://arxiv.org/abs/2304.13861