DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Morris, John X., Campion, Thomas R., Nutheti, Sri Laasya, Peng, Yifan, Raj, Akhil, Zabih, Ramin, Cole, Curtis L.
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916448891305984
author Morris, John X.
Campion, Thomas R.
Nutheti, Sri Laasya
Peng, Yifan
Raj, Akhil
Zabih, Ramin
Cole, Curtis L.
author_facet Morris, John X.
Campion, Thomas R.
Nutheti, Sri Laasya
Peng, Yifan
Raj, Akhil
Zabih, Ramin
Cole, Curtis L.
contents Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification methods are evaluated on highly saturated datasets (tools achieve near-perfect accuracy) which may not reflect the full variability or complexity of real-world clinical text and annotating them is resource intensive, which is a barrier to real-world applications. To address this gap, we developed an adversarial approach using a large language model (LLM) to re-identify the patient corresponding to a redacted clinical note and evaluated the performance with a novel De-Identification/Re-Identification (DIRI) method. Our method uses a large language model to reidentify the patient corresponding to a redacted clinical note. We demonstrate our method on medical data from Weill Cornell Medicine anonymized with three deidentification tools: rule-based Philter and two deep-learning-based models, BiLSTM-CRF and ClinicalBERT. Although ClinicalBERT was the most effective, masking all identified PII, our tool still reidentified 9% of clinical notes Our study highlights significant weaknesses in current deidentification technologies while providing a tool for iterative development and improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2410_17035
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
Morris, John X.
Campion, Thomas R.
Nutheti, Sri Laasya
Peng, Yifan
Raj, Akhil
Zabih, Ramin
Cole, Curtis L.
Computation and Language
Sharing protected health information (PHI) is critical for furthering biomedical research. Before data can be distributed, practitioners often perform deidentification to remove any PHI contained in the text. Contemporary deidentification methods are evaluated on highly saturated datasets (tools achieve near-perfect accuracy) which may not reflect the full variability or complexity of real-world clinical text and annotating them is resource intensive, which is a barrier to real-world applications. To address this gap, we developed an adversarial approach using a large language model (LLM) to re-identify the patient corresponding to a redacted clinical note and evaluated the performance with a novel De-Identification/Re-Identification (DIRI) method. Our method uses a large language model to reidentify the patient corresponding to a redacted clinical note. We demonstrate our method on medical data from Weill Cornell Medicine anonymized with three deidentification tools: rule-based Philter and two deep-learning-based models, BiLSTM-CRF and ClinicalBERT. Although ClinicalBERT was the most effective, masking all identified PII, our tool still reidentified 9% of clinical notes Our study highlights significant weaknesses in current deidentification technologies while providing a tool for iterative development and improvement.
title DIRI: Adversarial Patient Reidentification with Large Language Models for Evaluating Clinical Text Anonymization
topic Computation and Language
url https://arxiv.org/abs/2410.17035