FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915802878312448 |
|---|---|
| author | Nagesh, Nitish Shakibhamedan, Salar Bagheri, Mahdi Wang, Ziyu TaheriNejad, Nima Jantsch, Axel Rahmani, Amir M. |
| author_facet | Nagesh, Nitish Shakibhamedan, Salar Bagheri, Mahdi Wang, Ziyu TaheriNejad, Nima Jantsch, Axel Rahmani, Amir M. |
| contents | Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge about training generative models and require high computational resources. In this paper, we propose FairTabGen, an LLM-based tabular data generation framework that produces high-quality synthetic healthcare data using only a small subset of the original dataset. Our method combines in-context learning, prompt curation and embedding structural constraints for data synthesis. We evaluate performance on MIMIC-IV dataset. Our method using 99% less data and achieving 50% improvement for fairness through unawareness while maintaining competitive predictive utility. However, we observe data distribution of racial groups is skewed affecting demographic parity. We thereafter apply bias mitigation algorithms in the pre-processing stage, improving overall fairness by 10% highlighting effectiveness of our approach. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2508_11810 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples Nagesh, Nitish Shakibhamedan, Salar Bagheri, Mahdi Wang, Ziyu TaheriNejad, Nima Jantsch, Axel Rahmani, Amir M. Machine Learning Artificial Intelligence Synthetic healthcare data generation offers a promising solution to research limitations in clinical settings caused by privacy and regulatory constraints. However, current synthetic data generation approaches require specialized knowledge about training generative models and require high computational resources. In this paper, we propose FairTabGen, an LLM-based tabular data generation framework that produces high-quality synthetic healthcare data using only a small subset of the original dataset. Our method combines in-context learning, prompt curation and embedding structural constraints for data synthesis. We evaluate performance on MIMIC-IV dataset. Our method using 99% less data and achieving 50% improvement for fairness through unawareness while maintaining competitive predictive utility. However, we observe data distribution of racial groups is skewed affecting demographic parity. We thereafter apply bias mitigation algorithms in the pre-processing stage, improving overall fairness by 10% highlighting effectiveness of our approach. |
| title | FairTabGen: High-Fidelity and Fair Synthetic Health Data Generation from Limited Samples |
| topic | Machine Learning Artificial Intelligence |
| url | https://arxiv.org/abs/2508.11810 |