Exploring Membership Inference Vulnerabilities in Clinical Large Language Models

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Nemecek, Alexander, Yun, Zebin, Rahmani, Zahra, Harel, Yaniv, Chaudhary, Vipin, Sharif, Mahmood, Ayday, Erman
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914106177486848
author Nemecek, Alexander
Yun, Zebin
Rahmani, Zahra
Harel, Yaniv
Chaudhary, Vipin
Sharif, Mahmood
Ayday, Erman
author_facet Nemecek, Alexander
Yun, Zebin
Rahmani, Zahra
Harel, Yaniv
Chaudhary, Vipin
Sharif, Mahmood
Ayday, Erman
contents As large language models (LLMs) become progressively more embedded in clinical decision-support, documentation, and patient-information systems, ensuring their privacy and trustworthiness has emerged as an imperative challenge for the healthcare sector. Fine-tuning LLMs on sensitive electronic health record (EHR) data improves domain alignment but also raises the risk of exposing patient information through model behaviors. In this work-in-progress, we present an exploratory empirical study on membership inference vulnerabilities in clinical LLMs, focusing on whether adversaries can infer if specific patient records were used during model training. Using a state-of-the-art clinical question-answering model, Llemr, we evaluate both canonical loss-based attacks and a domain-motivated paraphrasing-based perturbation strategy that more realistically reflects clinical adversarial conditions. Our preliminary findings reveal limited but measurable membership leakage, suggesting that current clinical LLMs provide partial resistance yet remain susceptible to subtle privacy risks that could undermine trust in clinical AI adoption. These results motivate continued development of context-aware, domain-specific privacy evaluations and defenses such as differential privacy fine-tuning and paraphrase-aware training, to strengthen the security and trustworthiness of healthcare AI systems.
format Preprint
id arxiv_https___arxiv_org_abs_2510_18674
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
Nemecek, Alexander
Yun, Zebin
Rahmani, Zahra
Harel, Yaniv
Chaudhary, Vipin
Sharif, Mahmood
Ayday, Erman
Cryptography and Security
Artificial Intelligence
As large language models (LLMs) become progressively more embedded in clinical decision-support, documentation, and patient-information systems, ensuring their privacy and trustworthiness has emerged as an imperative challenge for the healthcare sector. Fine-tuning LLMs on sensitive electronic health record (EHR) data improves domain alignment but also raises the risk of exposing patient information through model behaviors. In this work-in-progress, we present an exploratory empirical study on membership inference vulnerabilities in clinical LLMs, focusing on whether adversaries can infer if specific patient records were used during model training. Using a state-of-the-art clinical question-answering model, Llemr, we evaluate both canonical loss-based attacks and a domain-motivated paraphrasing-based perturbation strategy that more realistically reflects clinical adversarial conditions. Our preliminary findings reveal limited but measurable membership leakage, suggesting that current clinical LLMs provide partial resistance yet remain susceptible to subtle privacy risks that could undermine trust in clinical AI adoption. These results motivate continued development of context-aware, domain-specific privacy evaluations and defenses such as differential privacy fine-tuning and paraphrase-aware training, to strengthen the security and trustworthiness of healthcare AI systems.
title Exploring Membership Inference Vulnerabilities in Clinical Large Language Models
topic Cryptography and Security
Artificial Intelligence
url https://arxiv.org/abs/2510.18674