Fairness-Optimized Synthetic EHR Generation for Arbitrary Downstream Predictive Tasks

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Tarek, Mirza Farhan Bin, Poulain, Raphael, Beheshti, Rahmatollah
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908425605087232
author Tarek, Mirza Farhan Bin
Poulain, Raphael
Beheshti, Rahmatollah
author_facet Tarek, Mirza Farhan Bin
Poulain, Raphael
Beheshti, Rahmatollah
contents Among various aspects of ensuring the responsible design of AI tools for healthcare applications, addressing fairness concerns has been a key focus area. Specifically, given the wide spread of electronic health record (EHR) data and their huge potential to inform a wide range of clinical decision support tasks, improving fairness in this category of health AI tools is of key importance. While such a broad problem (mitigating fairness in EHR-based AI models) has been tackled using various methods, task- and model-agnostic methods are noticeably rare. In this study, we aimed to target this gap by presenting a new pipeline that generates synthetic EHR data, which is not only consistent with (faithful to) the real EHR data but also can reduce the fairness concerns (defined by the end-user) in the downstream tasks, when combined with the real data. We demonstrate the effectiveness of our proposed pipeline across various downstream tasks and two different EHR datasets. Our proposed pipeline can add a widely applicable and complementary tool to the existing toolbox of methods to address fairness in health AI applications, such as those modifying the design of a downstream model. The codebase for our project is available at https://github.com/healthylaife/FairSynth
format Preprint
id arxiv_https___arxiv_org_abs_2406_02510
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fairness-Optimized Synthetic EHR Generation for Arbitrary Downstream Predictive Tasks
Tarek, Mirza Farhan Bin
Poulain, Raphael
Beheshti, Rahmatollah
Machine Learning
Among various aspects of ensuring the responsible design of AI tools for healthcare applications, addressing fairness concerns has been a key focus area. Specifically, given the wide spread of electronic health record (EHR) data and their huge potential to inform a wide range of clinical decision support tasks, improving fairness in this category of health AI tools is of key importance. While such a broad problem (mitigating fairness in EHR-based AI models) has been tackled using various methods, task- and model-agnostic methods are noticeably rare. In this study, we aimed to target this gap by presenting a new pipeline that generates synthetic EHR data, which is not only consistent with (faithful to) the real EHR data but also can reduce the fairness concerns (defined by the end-user) in the downstream tasks, when combined with the real data. We demonstrate the effectiveness of our proposed pipeline across various downstream tasks and two different EHR datasets. Our proposed pipeline can add a widely applicable and complementary tool to the existing toolbox of methods to address fairness in health AI applications, such as those modifying the design of a downstream model. The codebase for our project is available at https://github.com/healthylaife/FairSynth
title Fairness-Optimized Synthetic EHR Generation for Arbitrary Downstream Predictive Tasks
topic Machine Learning
url https://arxiv.org/abs/2406.02510