Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Ding, Bosheng, Qin, Chengwei, Zhao, Ruochen, Luo, Tianze, Li, Xinze, Chen, Guizhen, Xia, Wenhan, Hu, Junjie, Luu, Anh Tuan, Joty, Shafiq
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866929405298737152
author Ding, Bosheng
Qin, Chengwei
Zhao, Ruochen
Luo, Tianze
Li, Xinze
Chen, Guizhen
Xia, Wenhan
Hu, Junjie
Luu, Anh Tuan
Joty, Shafiq
author_facet Ding, Bosheng
Qin, Chengwei
Zhao, Ruochen
Luo, Tianze
Li, Xinze
Chen, Guizhen
Xia, Wenhan
Hu, Junjie
Luu, Anh Tuan
Joty, Shafiq
contents In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional data collection. This survey explores the transformative impact of LLMs on DA, particularly addressing the unique challenges and opportunities they present in the context of natural language processing (NLP) and beyond. From both data and learning perspectives, we examine various strategies that utilize LLMs for data augmentation, including a novel exploration of learning paradigms where LLM-generated data is used for diverse forms of further training. Additionally, this paper highlights the primary open challenges faced in this domain, ranging from controllable data augmentation to multi-modal data augmentation. This survey highlights a paradigm shift introduced by LLMs in DA, and aims to serve as a comprehensive guide for researchers and practitioners.
format Preprint
id arxiv_https___arxiv_org_abs_2403_02990
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
Ding, Bosheng
Qin, Chengwei
Zhao, Ruochen
Luo, Tianze
Li, Xinze
Chen, Guizhen
Xia, Wenhan
Hu, Junjie
Luu, Anh Tuan
Joty, Shafiq
Computation and Language
Artificial Intelligence
In the rapidly evolving field of large language models (LLMs), data augmentation (DA) has emerged as a pivotal technique for enhancing model performance by diversifying training examples without the need for additional data collection. This survey explores the transformative impact of LLMs on DA, particularly addressing the unique challenges and opportunities they present in the context of natural language processing (NLP) and beyond. From both data and learning perspectives, we examine various strategies that utilize LLMs for data augmentation, including a novel exploration of learning paradigms where LLM-generated data is used for diverse forms of further training. Additionally, this paper highlights the primary open challenges faced in this domain, ranging from controllable data augmentation to multi-modal data augmentation. This survey highlights a paradigm shift introduced by LLMs in DA, and aims to serve as a comprehensive guide for researchers and practitioners.
title Data Augmentation using Large Language Models: Data Perspectives, Learning Paradigms and Challenges
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2403.02990