An Investigation of Memorization Risk in Healthcare Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tonekaboni, Sana, Stempfle, Lena, Fallahpour, Adibvafa, Gerych, Walter, Ghassemi, Marzyeh
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908593664557056
author Tonekaboni, Sana
Stempfle, Lena
Fallahpour, Adibvafa
Gerych, Walter
Ghassemi, Marzyeh
author_facet Tonekaboni, Sana
Stempfle, Lena
Fallahpour, Adibvafa
Gerych, Walter
Ghassemi, Marzyeh
contents Foundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. In this work, we introduce a suite of black-box evaluation tests to assess privacy-related memorization risks in foundation models trained on structured EHR data. Our framework includes methods for probing memorization at both the embedding and generative levels, and aims to distinguish between model generalization and harmful memorization in clinically relevant settings. We contextualize memorization in terms of its potential to compromise patient privacy, particularly for vulnerable subgroups. We validate our approach on a publicly available EHR foundation model and release an open-source toolkit to facilitate reproducible and collaborative privacy assessments in healthcare AI.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12950
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Investigation of Memorization Risk in Healthcare Foundation Models
Tonekaboni, Sana
Stempfle, Lena
Fallahpour, Adibvafa
Gerych, Walter
Ghassemi, Marzyeh
Machine Learning
Foundation models trained on large-scale de-identified electronic health records (EHRs) hold promise for clinical applications. However, their capacity to memorize patient information raises important privacy concerns. In this work, we introduce a suite of black-box evaluation tests to assess privacy-related memorization risks in foundation models trained on structured EHR data. Our framework includes methods for probing memorization at both the embedding and generative levels, and aims to distinguish between model generalization and harmful memorization in clinically relevant settings. We contextualize memorization in terms of its potential to compromise patient privacy, particularly for vulnerable subgroups. We validate our approach on a publicly available EHR foundation model and release an open-source toolkit to facilitate reproducible and collaborative privacy assessments in healthcare AI.
title An Investigation of Memorization Risk in Healthcare Foundation Models
topic Machine Learning
url https://arxiv.org/abs/2510.12950