FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Nagesh, Nitish, Shakibhamedan, Salar, Bagheri, Mahdi, Wang, Ziyu, TaheriNejad, Nima, Jantsch, Axel, Rahmani, Amir M.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915802878312448
author Nagesh, Nitish
Shakibhamedan, Salar
Bagheri, Mahdi
Wang, Ziyu
TaheriNejad, Nima
Jantsch, Axel
Rahmani, Amir M.
author_facet Nagesh, Nitish
Shakibhamedan, Salar
Bagheri, Mahdi
Wang, Ziyu
TaheriNejad, Nima
Jantsch, Axel
Rahmani, Amir M.
contents Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge about training generative models and require high computational resources. In this paper, we propose FairTabGen, an LLM-based tabular data generation framework that produces high-quality synthetic healthcare data using only a small subset of the original dataset. Our method combines in-context learning, prompt curation and embedding structural constraints for data synthesis. We evaluate performance on MIMIC-IV dataset. Our method using 99% less data and achieving 50% improvement for fairness through unawareness while maintaining competitive predictive utility. However, we observe data distribution of racial groups is skewed affecting demographic parity. We thereafter apply bias mitigation algorithms in the pre-processing stage, improving overall fairness by 10% highlighting effectiveness of our approach.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11810
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples
Nagesh, Nitish
Shakibhamedan, Salar
Bagheri, Mahdi
Wang, Ziyu
TaheriNejad, Nima
Jantsch, Axel
Rahmani, Amir M.
Machine Learning
Artificial Intelligence
Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge about training generative models and require high computational resources. In this paper, we propose FairTabGen, an LLM-based tabular data generation framework that produces high-quality synthetic healthcare data using only a small subset of the original dataset. Our method combines in-context learning, prompt curation and embedding structural constraints for data synthesis. We evaluate performance on MIMIC-IV dataset. Our method using 99% less data and achieving 50% improvement for fairness through unawareness while maintaining competitive predictive utility. However, we observe data distribution of racial groups is skewed affecting demographic parity. We thereafter apply bias mitigation algorithms in the pre-processing stage, improving overall fairness by 10% highlighting effectiveness of our approach.
title FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.11810