NFDI4Health workflow and service for synthetic data generation, assessment and risk management

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Moazemi, Sobhan, Adams, Tim, NG, Hwei Geok, Kühnel, Lisa, Schneider, Julian, Näher, Anatol-Fiete, Fluck, Juliane, Fröhlich, Holger
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866911982316158976
author Moazemi, Sobhan
Adams, Tim
NG, Hwei Geok
Kühnel, Lisa
Schneider, Julian
Näher, Anatol-Fiete
Fluck, Juliane
Fröhlich, Holger
author_facet Moazemi, Sobhan
Adams, Tim
NG, Hwei Geok
Kühnel, Lisa
Schneider, Julian
Näher, Anatol-Fiete
Fluck, Juliane
Fröhlich, Holger
contents Individual health data is crucial for scientific advancements, particularly in developing Artificial Intelligence (AI); however, sharing real patient information is often restricted due to privacy concerns. A promising solution to this challenge is synthetic data generation. This technique creates entirely new datasets that mimic the statistical properties of real data, while preserving confidential patient information. In this paper, we present the workflow and different services developed in the context of Germany's National Data Infrastructure project NFDI4Health. First, two state-of-the-art AI tools (namely, VAMBN and MultiNODEs) for generating synthetic health data are outlined. Further, we introduce SYNDAT (a public web-based tool) which allows users to visualize and assess the quality and risk of synthetic data provided by desired generative models. Additionally, the utility of the proposed methods and the web-based tool is showcased using data from Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Center for Cancer Registry Data of the Robert Koch Institute (RKI).
format Preprint
id arxiv_https___arxiv_org_abs_2408_04478
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle NFDI4Health workflow and service for synthetic data generation, assessment and risk management
Moazemi, Sobhan
Adams, Tim
NG, Hwei Geok
Kühnel, Lisa
Schneider, Julian
Näher, Anatol-Fiete
Fluck, Juliane
Fröhlich, Holger
Machine Learning
Individual health data is crucial for scientific advancements, particularly in developing Artificial Intelligence (AI); however, sharing real patient information is often restricted due to privacy concerns. A promising solution to this challenge is synthetic data generation. This technique creates entirely new datasets that mimic the statistical properties of real data, while preserving confidential patient information. In this paper, we present the workflow and different services developed in the context of Germany's National Data Infrastructure project NFDI4Health. First, two state-of-the-art AI tools (namely, VAMBN and MultiNODEs) for generating synthetic health data are outlined. Further, we introduce SYNDAT (a public web-based tool) which allows users to visualize and assess the quality and risk of synthetic data provided by desired generative models. Additionally, the utility of the proposed methods and the web-based tool is showcased using data from Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Center for Cancer Registry Data of the Robert Koch Institute (RKI).
title NFDI4Health workflow and service for synthetic data generation, assessment and risk management
topic Machine Learning
url https://arxiv.org/abs/2408.04478