TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Sheng, Chen, Pengan, Zhou, Jingqi, Li, Qintong, Dong, Jingwei, Gao, Jiahui, Xue, Boyang, Jiang, Jiyue, Kong, Lingpeng, Wu, Chuan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918067195346944
author Wang, Sheng
Chen, Pengan
Zhou, Jingqi
Li, Qintong
Dong, Jingwei
Gao, Jiahui
Xue, Boyang
Jiang, Jiyue
Kong, Lingpeng
Wu, Chuan
author_facet Wang, Sheng
Chen, Pengan
Zhou, Jingqi
Li, Qintong
Dong, Jingwei
Gao, Jiahui
Xue, Boyang
Jiang, Jiyue
Kong, Lingpeng
Wu, Chuan
contents Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases, and low-variation prompts, resulting in limited diversity and biased distributions with the increase of data scales. To tackle this challenge, we introduce TREESYNTH, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness before synthesizing samples within each atomic subspace. This globally dividing-and-synthesizing method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the rebalancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently demonstrate the superior data diversity, model performance, and robust scalability of TREESYNTH compared to both human-crafted datasets and peer data synthesis methods, with an average performance gain reaching 10%. Besides, the consistent improvements of TREESYNTH-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth.
format Preprint
id arxiv_https___arxiv_org_abs_2503_17195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning
Wang, Sheng
Chen, Pengan
Zhou, Jingqi
Li, Qintong
Dong, Jingwei
Gao, Jiahui
Xue, Boyang
Jiang, Jiyue
Kong, Lingpeng
Wu, Chuan
Machine Learning
Artificial Intelligence
Model customization necessitates high-quality and diverse datasets, but acquiring such data remains time-consuming and labor-intensive. Despite the great potential of large language models (LLMs) for data synthesis, current approaches are constrained by limited seed data, model biases, and low-variation prompts, resulting in limited diversity and biased distributions with the increase of data scales. To tackle this challenge, we introduce TREESYNTH, a tree-guided subspace-based data synthesis approach inspired by decision trees. It constructs a spatial partitioning tree to recursively divide a task-specific full data space (i.e., root node) into numerous atomic subspaces (i.e., leaf nodes) with mutually exclusive and exhaustive attributes to ensure both distinctiveness and comprehensiveness before synthesizing samples within each atomic subspace. This globally dividing-and-synthesizing method finally collects subspace samples into a comprehensive dataset, effectively circumventing repetition and space collapse to ensure the diversity of large-scale data synthesis. Furthermore, the spatial partitioning tree enables sample allocation into atomic subspaces, allowing the rebalancing of existing datasets for more balanced and comprehensive distributions. Empirically, extensive experiments across diverse benchmarks consistently demonstrate the superior data diversity, model performance, and robust scalability of TREESYNTH compared to both human-crafted datasets and peer data synthesis methods, with an average performance gain reaching 10%. Besides, the consistent improvements of TREESYNTH-balanced datasets highlight its efficacious application to redistribute existing datasets for more comprehensive coverage and the induced performance enhancement. The code is available at https://github.com/cpa2001/TreeSynth.
title TreeSynth: Synthesizing Diverse Data from Scratch via Tree-Guided Subspace Partitioning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.17195