Data Generation Using Large Language Models for Text Classification: An Empirical Case Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yinheng, Bonatti, Rogerio, Abdali, Sara, Wagle, Justin, Koishida, Kazuhito
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916331228495872
author Li, Yinheng
Bonatti, Rogerio
Abdali, Sara
Wagle, Justin
Koishida, Kazuhito
author_facet Li, Yinheng
Bonatti, Rogerio
Abdali, Sara
Wagle, Justin
Koishida, Kazuhito
contents Using Large Language Models (LLMs) to generate synthetic data for model training has become increasingly popular in recent years. While LLMs are capable of producing realistic training data, the effectiveness of data generation is influenced by various factors, including the choice of prompt, task complexity, and the quality, quantity, and diversity of the generated data. In this work, we focus exclusively on using synthetic data for text classification tasks. Specifically, we use natural language understanding (NLU) models trained on synthetic data to assess the quality of synthetic data from different generation approaches. This work provides an empirical analysis of the impact of these factors and offers recommendations for better data generation practices.
format Preprint
id arxiv_https___arxiv_org_abs_2407_12813
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
Li, Yinheng
Bonatti, Rogerio
Abdali, Sara
Wagle, Justin
Koishida, Kazuhito
Computation and Language
Artificial Intelligence
Using Large Language Models (LLMs) to generate synthetic data for model training has become increasingly popular in recent years. While LLMs are capable of producing realistic training data, the effectiveness of data generation is influenced by various factors, including the choice of prompt, task complexity, and the quality, quantity, and diversity of the generated data. In this work, we focus exclusively on using synthetic data for text classification tasks. Specifically, we use natural language understanding (NLU) models trained on synthetic data to assess the quality of synthetic data from different generation approaches. This work provides an empirical analysis of the impact of these factors and offers recommendations for better data generation practices.
title Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2407.12813