SEDGE: Structural Extrapolated Data Generation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Kun, Sun, Jiaqi, Li, Yiqing, Ng, Ignavier, Deka, Namrata, Xie, Shaoan
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866913125323767808
author Zhang, Kun
Sun, Jiaqi
Li, Yiqing
Ng, Ignavier
Deka, Namrata
Xie, Shaoan
author_facet Zhang, Kun
Sun, Jiaqi
Li, Yiqing
Ng, Ignavier
Deka, Namrata
Xie, Shaoan
contents This paper aims to address the challenge of data generation beyond the training data and proposes a framework for Structural Extrapolated Data GEneration (SEDGE) based on suitable assumptions on the underlying data-generating process. We provide conditions under which data satisfying novel specifications can be generated reliably, together with the approximate identifiability of the distribution of such data under certain ``conservative" assumptions, as well as the inherent non-identifiability of this distribution without such assumptions. On the algorithmic side, we develop practical methods to achieve extrapolated data generation, based on a structure-informed optimization strategy or diffusion posterior sampling, respectively. We verify the extrapolation performance on synthetic data and also consider extrapolated image generation as a real-world scenario to illustrate the validity of the proposed framework.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02482
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle SEDGE: Structural Extrapolated Data Generation
Zhang, Kun
Sun, Jiaqi
Li, Yiqing
Ng, Ignavier
Deka, Namrata
Xie, Shaoan
Machine Learning
This paper aims to address the challenge of data generation beyond the training data and proposes a framework for Structural Extrapolated Data GEneration (SEDGE) based on suitable assumptions on the underlying data-generating process. We provide conditions under which data satisfying novel specifications can be generated reliably, together with the approximate identifiability of the distribution of such data under certain ``conservative" assumptions, as well as the inherent non-identifiability of this distribution without such assumptions. On the algorithmic side, we develop practical methods to achieve extrapolated data generation, based on a structure-informed optimization strategy or diffusion posterior sampling, respectively. We verify the extrapolation performance on synthetic data and also consider extrapolated image generation as a real-world scenario to illustrate the validity of the proposed framework.
title SEDGE: Structural Extrapolated Data Generation
topic Machine Learning
url https://arxiv.org/abs/2604.02482