DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Ying, Wang, Xinyao, Niu, Yulei, Shen, Yaojie, Tang, Lexin, Chen, Fan, He, Ben, Sun, Le, Wen, Longyin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908401202626560
author Zhou, Ying
Wang, Xinyao
Niu, Yulei
Shen, Yaojie
Tang, Lexin
Chen, Fan
He, Ben
Sun, Le
Wen, Longyin
author_facet Zhou, Ying
Wang, Xinyao
Niu, Yulei
Shen, Yaojie
Tang, Lexin
Chen, Fan
He, Ben
Sun, Le
Wen, Longyin
contents Recent advancements in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis. However, synthetic data generation via prompting LLMs remains challenging due to LLMs' limited understanding of target data distributions and the complexity of prompt engineering, especially for structured formatted data. To address these issues, we introduce DiffLM, a controllable data synthesis framework based on variational autoencoder (VAE), which further (1) leverages diffusion models to reserve more information of original distribution and format structure in the learned latent distribution and (2) decouples the learning of target distribution knowledge from the LLM's generative objectives via a plug-and-play latent feature injection module. As we observed significant discrepancies between the VAE's latent representations and the real data distribution, the latent diffusion module is introduced into our framework to learn a fully expressive latent distribution. Evaluations on seven real-world datasets with structured formatted data (i.e., Tabular, Code, and Tool data) demonstrate that DiffLM generates high-quality data, with performance on downstream tasks surpassing that of real data by 2%-7% in certain cases. Data and code are available at https://github.com/bytedance/DiffLM.
format Preprint
id arxiv_https___arxiv_org_abs_2411_03250
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models
Zhou, Ying
Wang, Xinyao
Niu, Yulei
Shen, Yaojie
Tang, Lexin
Chen, Fan
He, Ben
Sun, Le
Wen, Longyin
Machine Learning
Artificial Intelligence
Computation and Language
Recent advancements in large language models (LLMs) have significantly enhanced their knowledge and generative capabilities, leading to a surge of interest in leveraging LLMs for high-quality data synthesis. However, synthetic data generation via prompting LLMs remains challenging due to LLMs' limited understanding of target data distributions and the complexity of prompt engineering, especially for structured formatted data. To address these issues, we introduce DiffLM, a controllable data synthesis framework based on variational autoencoder (VAE), which further (1) leverages diffusion models to reserve more information of original distribution and format structure in the learned latent distribution and (2) decouples the learning of target distribution knowledge from the LLM's generative objectives via a plug-and-play latent feature injection module. As we observed significant discrepancies between the VAE's latent representations and the real data distribution, the latent diffusion module is introduced into our framework to learn a fully expressive latent distribution. Evaluations on seven real-world datasets with structured formatted data (i.e., Tabular, Code, and Tool data) demonstrate that DiffLM generates high-quality data, with performance on downstream tasks surpassing that of real data by 2%-7% in certain cases. Data and code are available at https://github.com/bytedance/DiffLM.
title DiffLM: Controllable Synthetic Data Generation via Diffusion Language Models
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.03250