Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Choi, Juhwan, Yun, Jungmin, Jin, Kyohoon, Kim, YoungBin
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910617743392768
author Choi, Juhwan
Yun, Jungmin
Jin, Kyohoon
Kim, YoungBin
author_facet Choi, Juhwan
Yun, Jungmin
Jin, Kyohoon
Kim, YoungBin
contents The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models. However, datasets often contain noisy data inadvertently included during the construction process. Numerous attempts have been made to correct this issue through human annotators. However, hiring and managing human annotators is expensive and time-consuming. As an alternative, recent studies are exploring the use of large language models (LLMs) for data annotation. In this study, we present a case study that extends the application of LLM-based data annotation to enhance the quality of existing datasets through a cleansing strategy. Specifically, we leverage approaches such as chain-of-thought and majority voting to imitate human annotation and classify unrelated documents from the Multi-News dataset, which is widely used for the multi-document summarization task. Through our proposed cleansing method, we introduce an enhanced Multi-News+. By employing LLMs for data cleansing, we demonstrate an efficient and effective approach to improving dataset quality without relying on expensive human annotation efforts.
format Preprint
id arxiv_https___arxiv_org_abs_2404_09682
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
Choi, Juhwan
Yun, Jungmin
Jin, Kyohoon
Kim, YoungBin
Computation and Language
Artificial Intelligence
The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models. However, datasets often contain noisy data inadvertently included during the construction process. Numerous attempts have been made to correct this issue through human annotators. However, hiring and managing human annotators is expensive and time-consuming. As an alternative, recent studies are exploring the use of large language models (LLMs) for data annotation. In this study, we present a case study that extends the application of LLM-based data annotation to enhance the quality of existing datasets through a cleansing strategy. Specifically, we leverage approaches such as chain-of-thought and majority voting to imitate human annotation and classify unrelated documents from the Multi-News dataset, which is widely used for the multi-document summarization task. Through our proposed cleansing method, we introduce an enhanced Multi-News+. By employing LLMs for data cleansing, we demonstrate an efficient and effective approach to improving dataset quality without relying on expensive human annotation efforts.
title Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.09682