DataGen: Unified Synthetic Dataset Generation via Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Yue, Wu, Siyuan, Gao, Chujie, Chen, Dongping, Zhang, Qihui, Wan, Yao, Zhou, Tianyi, Gao, Jianfeng, Xiao, Chaowei, Sun, Lichao, Zhang, Xiangliang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908657929682944
author Huang, Yue
Wu, Siyuan
Gao, Chujie
Chen, Dongping
Zhang, Qihui
Wan, Yao
Zhou, Tianyi
Gao, Jianfeng
Xiao, Chaowei
Sun, Lichao
Zhang, Xiangliang
author_facet Huang, Yue
Wu, Siyuan
Gao, Chujie
Chen, Dongping
Zhang, Qihui
Wan, Yao
Zhou, Tianyi
Gao, Jianfeng
Xiao, Chaowei
Sun, Lichao
Zhang, Xiangliang
contents Large Language Models (LLMs) such as GPT-4 and Llama3 have significantly impacted various fields by enabling high-quality synthetic data generation and reducing dependence on expensive human-generated datasets. Despite this, challenges remain in the areas of generalization, controllability, diversity, and truthfulness within the existing generative frameworks. To address these challenges, this paper presents DataGen, a comprehensive LLM-powered framework designed to produce diverse, accurate, and highly controllable datasets. DataGen is adaptable, supporting all types of text datasets and enhancing the generative process through innovative mechanisms. To augment data diversity, DataGen incorporates an attribute-guided generation module and a group checking feature. For accuracy, it employs a code-based mathematical assessment for label verification alongside a retrieval-augmented generation technique for factual validation. The framework also allows for user-specified constraints, enabling customization of the data generation process to suit particular requirements. Extensive experiments demonstrate the superior quality of data generated by DataGen, and each module within DataGen plays a critical role in this enhancement. Additionally, DataGen is applied in two practical scenarios: benchmarking LLMs and data augmentation. The results indicate that DataGen effectively supports dynamic and evolving benchmarking and that data augmentation improves LLM capabilities in various domains, including agent-oriented abilities and reasoning skills.
format Preprint
id arxiv_https___arxiv_org_abs_2406_18966
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DataGen: Unified Synthetic Dataset Generation via Large Language Models
Huang, Yue
Wu, Siyuan
Gao, Chujie
Chen, Dongping
Zhang, Qihui
Wan, Yao
Zhou, Tianyi
Gao, Jianfeng
Xiao, Chaowei
Sun, Lichao
Zhang, Xiangliang
Computation and Language
Large Language Models (LLMs) such as GPT-4 and Llama3 have significantly impacted various fields by enabling high-quality synthetic data generation and reducing dependence on expensive human-generated datasets. Despite this, challenges remain in the areas of generalization, controllability, diversity, and truthfulness within the existing generative frameworks. To address these challenges, this paper presents DataGen, a comprehensive LLM-powered framework designed to produce diverse, accurate, and highly controllable datasets. DataGen is adaptable, supporting all types of text datasets and enhancing the generative process through innovative mechanisms. To augment data diversity, DataGen incorporates an attribute-guided generation module and a group checking feature. For accuracy, it employs a code-based mathematical assessment for label verification alongside a retrieval-augmented generation technique for factual validation. The framework also allows for user-specified constraints, enabling customization of the data generation process to suit particular requirements. Extensive experiments demonstrate the superior quality of data generated by DataGen, and each module within DataGen plays a critical role in this enhancement. Additionally, DataGen is applied in two practical scenarios: benchmarking LLMs and data augmentation. The results indicate that DataGen effectively supports dynamic and evolving benchmarking and that data augmentation improves LLM capabilities in various domains, including agent-oriented abilities and reasoning skills.
title DataGen: Unified Synthetic Dataset Generation via Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2406.18966