Opinion: Revisiting synthetic data classifications from a privacy perspective

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Vallevik, Vibeke Binz, Marshall, Serena Elizabeth, Babic, Aleksandar, Nygaard, Jan Franz
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915431320649728
author Vallevik, Vibeke Binz
Marshall, Serena Elizabeth
Babic, Aleksandar
Nygaard, Jan Franz
author_facet Vallevik, Vibeke Binz
Marshall, Serena Elizabeth
Babic, Aleksandar
Nygaard, Jan Franz
contents Synthetic data is emerging as a cost-effective solution necessary to meet the increasing data demands of AI development, created either from existing knowledge or derived from real data. The traditional classification of synthetic data types into hybrid, partial or fully synthetic datasets has limited value and does not reflect the ever-increasing methods to generate synthetic data. The generation method and their source jointly shape the characteristics of synthetic data, which in turn determines its practical applications. We make a case for an alternative approach to grouping synthetic data types that better reflect privacy perspectives in order to facilitate regulatory guidance in the generation and processing of synthetic data. This approach to classification provides flexibility to new advancements like deep generative methods and offers a more practical framework for future applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_03506
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Opinion: Revisiting synthetic data classifications from a privacy perspective
Vallevik, Vibeke Binz
Marshall, Serena Elizabeth
Babic, Aleksandar
Nygaard, Jan Franz
Machine Learning
Artificial Intelligence
Synthetic data is emerging as a cost-effective solution necessary to meet the increasing data demands of AI development, created either from existing knowledge or derived from real data. The traditional classification of synthetic data types into hybrid, partial or fully synthetic datasets has limited value and does not reflect the ever-increasing methods to generate synthetic data. The generation method and their source jointly shape the characteristics of synthetic data, which in turn determines its practical applications. We make a case for an alternative approach to grouping synthetic data types that better reflect privacy perspectives in order to facilitate regulatory guidance in the generation and processing of synthetic data. This approach to classification provides flexibility to new advancements like deep generative methods and offers a more practical framework for future applications.
title Opinion: Revisiting synthetic data classifications from a privacy perspective
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2503.03506