Synthetic Tabular Data: Methods, Attacks and Defenses
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908396485083136 |
|---|---|
| author | Cormode, Graham Maddock, Samuel Ullah, Enayat Gade, Shripad |
| author_facet | Cormode, Graham Maddock, Samuel Ullah, Enayat Gade, Shripad |
| contents | Synthetic data is often positioned as a solution to replace sensitive fixed-size datasets with a source of unlimited matching data, freed from privacy concerns. There has been much progress in synthetic data generation over the last decade, leveraging corresponding advances in machine learning and data analytics. In this survey, we cover the key developments and the main concepts in tabular synthetic data generation, including paradigms based on probabilistic graphical models and on deep learning. We provide background and motivation, before giving a technical deep-dive into the methodologies. We also address the limitations of synthetic data, by studying attacks that seek to retrieve information about the original sensitive data. Finally, we present extensions and open problems in this area. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_06108 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Synthetic Tabular Data: Methods, Attacks and Defenses Cormode, Graham Maddock, Samuel Ullah, Enayat Gade, Shripad Machine Learning Cryptography and Security Synthetic data is often positioned as a solution to replace sensitive fixed-size datasets with a source of unlimited matching data, freed from privacy concerns. There has been much progress in synthetic data generation over the last decade, leveraging corresponding advances in machine learning and data analytics. In this survey, we cover the key developments and the main concepts in tabular synthetic data generation, including paradigms based on probabilistic graphical models and on deep learning. We provide background and motivation, before giving a technical deep-dive into the methodologies. We also address the limitations of synthetic data, by studying attacks that seek to retrieve information about the original sensitive data. Finally, we present extensions and open problems in this area. |
| title | Synthetic Tabular Data: Methods, Attacks and Defenses |
| topic | Machine Learning Cryptography and Security |
| url | https://arxiv.org/abs/2506.06108 |