Empowering Large Language Models for Textual Data Augmentation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Yichuan, Ding, Kaize, Wang, Jianling, Lee, Kyumin
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866910426237763584
author Li, Yichuan
Ding, Kaize
Wang, Jianling
Lee, Kyumin
author_facet Li, Yichuan
Ding, Kaize
Wang, Jianling
Lee, Kyumin
contents With the capabilities of understanding and executing natural language instructions, Large language models (LLMs) can potentially act as a powerful tool for textual data augmentation. However, the quality of augmented data depends heavily on the augmentation instructions provided, and the effectiveness can fluctuate across different downstream tasks. While manually crafting and selecting instructions can offer some improvement, this approach faces scalability and consistency issues in practice due to the diversity of downstream tasks. In this work, we address these limitations by proposing a new solution, which can automatically generate a large pool of augmentation instructions and select the most suitable task-informed instructions, thereby empowering LLMs to create high-quality augmented data for different downstream tasks. Empirically, the proposed approach consistently generates augmented data with better quality compared to non-LLM and LLM-based data augmentation methods, leading to the best performance on 26 few-shot learning tasks sourced from a wide range of application domains.
format Preprint
id arxiv_https___arxiv_org_abs_2404_17642
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Empowering Large Language Models for Textual Data Augmentation
Li, Yichuan
Ding, Kaize
Wang, Jianling
Lee, Kyumin
Computation and Language
Artificial Intelligence
With the capabilities of understanding and executing natural language instructions, Large language models (LLMs) can potentially act as a powerful tool for textual data augmentation. However, the quality of augmented data depends heavily on the augmentation instructions provided, and the effectiveness can fluctuate across different downstream tasks. While manually crafting and selecting instructions can offer some improvement, this approach faces scalability and consistency issues in practice due to the diversity of downstream tasks. In this work, we address these limitations by proposing a new solution, which can automatically generate a large pool of augmentation instructions and select the most suitable task-informed instructions, thereby empowering LLMs to create high-quality augmented data for different downstream tasks. Empirically, the proposed approach consistently generates augmented data with better quality compared to non-LLM and LLM-based data augmentation methods, leading to the best performance on 26 few-shot learning tasks sourced from a wide range of application domains.
title Empowering Large Language Models for Textual Data Augmentation
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2404.17642