Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pissarra, David, Curioso, Isabel, Alveira, João, Pereira, Duarte, Ribeiro, Bruno, Souper, Tomás, Gomes, Vasco, Carreiro, André V., Rolla, Vitor
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929372854747136
author Pissarra, David
Curioso, Isabel
Alveira, João
Pereira, Duarte
Ribeiro, Bruno
Souper, Tomás
Gomes, Vasco
Carreiro, André V.
Rolla, Vitor
author_facet Pissarra, David
Curioso, Isabel
Alveira, João
Pereira, Duarte
Ribeiro, Bruno
Souper, Tomás
Gomes, Vasco
Carreiro, André V.
Rolla, Vitor
contents Automated clinical text anonymization has the potential to unlock the widespread sharing of textual health data for secondary usage while assuring patient privacy and safety. Despite the proposal of many complex and theoretically successful anonymization solutions in literature, these techniques remain flawed. As such, clinical institutions are still reluctant to apply them for open access to their data. Recent advances in developing Large Language Models (LLMs) pose a promising opportunity to further the field, given their capability to perform various tasks. This paper proposes six new evaluation metrics tailored to the challenges of generative anonymization with LLMs. Moreover, we present a comparative study of LLM-based methods, testing them against two baseline techniques. Our results establish LLM-based models as a reliable alternative to common approaches, paving the way toward trustworthy anonymization of clinical text.
format Preprint
id arxiv_https___arxiv_org_abs_2406_00062
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study
Pissarra, David
Curioso, Isabel
Alveira, João
Pereira, Duarte
Ribeiro, Bruno
Souper, Tomás
Gomes, Vasco
Carreiro, André V.
Rolla, Vitor
Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
I.2.7
Automated clinical text anonymization has the potential to unlock the widespread sharing of textual health data for secondary usage while assuring patient privacy and safety. Despite the proposal of many complex and theoretically successful anonymization solutions in literature, these techniques remain flawed. As such, clinical institutions are still reluctant to apply them for open access to their data. Recent advances in developing Large Language Models (LLMs) pose a promising opportunity to further the field, given their capability to perform various tasks. This paper proposes six new evaluation metrics tailored to the challenges of generative anonymization with LLMs. Moreover, we present a comparative study of LLM-based methods, testing them against two baseline techniques. Our results establish LLM-based models as a reliable alternative to common approaches, paving the way toward trustworthy anonymization of clinical text.
title Unlocking the Potential of Large Language Models for Clinical Text Anonymization: A Comparative Study
topic Computation and Language
Artificial Intelligence
Cryptography and Security
Machine Learning
I.2.7
url https://arxiv.org/abs/2406.00062