Reasoning-Driven Synthetic Data Generation and Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Davidson, Tim R., Seguin, Benoit, Bacis, Enrico, Ilharco, Cesar, Harkous, Hamza
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912992285687808
author Davidson, Tim R.
Seguin, Benoit
Bacis, Enrico
Ilharco, Cesar
Harkous, Hamza
author_facet Davidson, Tim R.
Seguin, Benoit
Bacis, Enrico
Ilharco, Cesar
Harkous, Hamza
contents Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and time-consuming, leading model builders to increasingly consider synthetic data as a scalable alternative. However, existing synthetic data generation methods often rely on manual prompts, evolutionary algorithms, or extensive seed data from the target distribution - limiting their scalability, explainability, and control. In this paper, we introduce Simula: a novel reasoning-driven framework for data generation and evaluation. It employs a seedless, agentic approach to generate synthetic datasets at scale, allowing users to define desired dataset characteristics through an explainable and controllable process that enables fine-grained resource allocation. We show the efficacy of our approach on a variety of datasets, rigorously testing both intrinsic and downstream properties. Our work (1) offers guidelines for synthetic data mechanism design, (2) provides insights into generating and evaluating synthetic data at scale, and (3) unlocks new opportunities for developing and deploying AI in domains where data scarcity or privacy concerns are paramount.
format Preprint
id arxiv_https___arxiv_org_abs_2603_29791
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Reasoning-Driven Synthetic Data Generation and Evaluation
Davidson, Tim R.
Seguin, Benoit
Bacis, Enrico
Ilharco, Cesar
Harkous, Hamza
Artificial Intelligence
Computation and Language
Machine Learning
Although many AI applications of interest require specialized multi-modal models, relevant data to train such models is inherently scarce or inaccessible. Filling these gaps with human annotators is prohibitively expensive, error-prone, and time-consuming, leading model builders to increasingly consider synthetic data as a scalable alternative. However, existing synthetic data generation methods often rely on manual prompts, evolutionary algorithms, or extensive seed data from the target distribution - limiting their scalability, explainability, and control. In this paper, we introduce Simula: a novel reasoning-driven framework for data generation and evaluation. It employs a seedless, agentic approach to generate synthetic datasets at scale, allowing users to define desired dataset characteristics through an explainable and controllable process that enables fine-grained resource allocation. We show the efficacy of our approach on a variety of datasets, rigorously testing both intrinsic and downstream properties. Our work (1) offers guidelines for synthetic data mechanism design, (2) provides insights into generating and evaluating synthetic data at scale, and (3) unlocks new opportunities for developing and deploying AI in domains where data scarcity or privacy concerns are paramount.
title Reasoning-Driven Synthetic Data Generation and Evaluation
topic Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2603.29791