Can large language models be privacy preserving and fair medical coders?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dadsetan, Ali, Soleymani, Dorsa, Zeng, Xijie, Rudzicz, Frank
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916512817741824
author Dadsetan, Ali
Soleymani, Dorsa
Zeng, Xijie
Rudzicz, Frank
author_facet Dadsetan, Ali
Soleymani, Dorsa
Zeng, Xijie
Rudzicz, Frank
contents Protecting patient data privacy is a critical concern when deploying machine learning algorithms in healthcare. Differential privacy (DP) is a common method for preserving privacy in such settings and, in this work, we examine two key trade-offs in applying DP to the NLP task of medical coding (ICD classification). Regarding the privacy-utility trade-off, we observe a significant performance drop in the privacy preserving models, with more than a 40% reduction in micro F1 scores on the top 50 labels in the MIMIC-III dataset. From the perspective of the privacy-fairness trade-off, we also observe an increase of over 3% in the recall gap between male and female patients in the DP models. Further understanding these trade-offs will help towards the challenges of real-world deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2412_05533
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Can large language models be privacy preserving and fair medical coders?
Dadsetan, Ali
Soleymani, Dorsa
Zeng, Xijie
Rudzicz, Frank
Machine Learning
Cryptography and Security
Protecting patient data privacy is a critical concern when deploying machine learning algorithms in healthcare. Differential privacy (DP) is a common method for preserving privacy in such settings and, in this work, we examine two key trade-offs in applying DP to the NLP task of medical coding (ICD classification). Regarding the privacy-utility trade-off, we observe a significant performance drop in the privacy preserving models, with more than a 40% reduction in micro F1 scores on the top 50 labels in the MIMIC-III dataset. From the perspective of the privacy-fairness trade-off, we also observe an increase of over 3% in the recall gap between male and female patients in the DP models. Further understanding these trade-offs will help towards the challenges of real-world deployment.
title Can large language models be privacy preserving and fair medical coders?
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2412.05533