Medical Image De-Identification Resources: Synthetic DICOM Data and Tools for Validation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rutherford, Michael W., Nolan, Tracy, Pei, Linmin, Wagner, Ulrike, Pan, Qinyan, Farmer, Phillip, Smith, Kirk, Kopchick, Benjamin, Opsahl-Ong, Laura, Sutton, Granger, Clunie, David, Farahani, Keyvan, Prior, Fred
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911088949329920
author Rutherford, Michael W.
Nolan, Tracy
Pei, Linmin
Wagner, Ulrike
Pan, Qinyan
Farmer, Phillip
Smith, Kirk
Kopchick, Benjamin
Opsahl-Ong, Laura
Sutton, Granger
Clunie, David
Farahani, Keyvan
Prior, Fred
author_facet Rutherford, Michael W.
Nolan, Tracy
Pei, Linmin
Wagner, Ulrike
Pan, Qinyan
Farmer, Phillip
Smith, Kirk
Kopchick, Benjamin
Opsahl-Ong, Laura
Sutton, Granger
Clunie, David
Farahani, Keyvan
Prior, Fred
contents Medical imaging research increasingly depends on large-scale data sharing to promote reproducibility and train Artificial Intelligence (AI) models. Ensuring patient privacy remains a significant challenge for open-access data sharing. Digital Imaging and Communications in Medicine (DICOM), the global standard data format for medical imaging, encodes both essential clinical metadata and extensive protected health information (PHI) and personally identifiable information (PII). Effective de-identification must remove identifiers, preserve scientific utility, and maintain DICOM validity. Tools exist to perform de-identification, but few assess its effectiveness, and most rely on subjective reviews, limiting reproducibility and regulatory confidence. To address this gap, we developed an openly accessible DICOM dataset infused with synthetic PHI/PII and an evaluation framework for benchmarking image de-identification workflows. The Medical Image de-identification (MIDI) dataset was built using publicly available de-identified data from The Cancer Imaging Archive (TCIA). It includes 538 subjects (216 for validation, 322 for testing), 605 studies, 708 series, and 53,581 DICOM image instances. These span multiple vendors, imaging modalities, and cancer types. Synthetic PHI and PII were embedded into structured data elements, plain text data elements, and pixel data to simulate real-world identity leaks encountered by TCIA curation teams. Accompanying evaluation tools include a Python script, answer keys (known truth), and mapping files that enable automated comparison of curated data against expected transformations. The framework is aligned with the HIPAA Privacy Rule "Safe Harbor" method, DICOM PS3.15 Confidentiality Profiles, and TCIA best practices. It supports objective, standards-driven evaluation of de-identification workflows, promoting safer and more consistent medical image sharing.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01889
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Medical Image De-Identification Resources: Synthetic DICOM Data and Tools for Validation
Rutherford, Michael W.
Nolan, Tracy
Pei, Linmin
Wagner, Ulrike
Pan, Qinyan
Farmer, Phillip
Smith, Kirk
Kopchick, Benjamin
Opsahl-Ong, Laura
Sutton, Granger
Clunie, David
Farahani, Keyvan
Prior, Fred
Computer Vision and Pattern Recognition
Medical imaging research increasingly depends on large-scale data sharing to promote reproducibility and train Artificial Intelligence (AI) models. Ensuring patient privacy remains a significant challenge for open-access data sharing. Digital Imaging and Communications in Medicine (DICOM), the global standard data format for medical imaging, encodes both essential clinical metadata and extensive protected health information (PHI) and personally identifiable information (PII). Effective de-identification must remove identifiers, preserve scientific utility, and maintain DICOM validity. Tools exist to perform de-identification, but few assess its effectiveness, and most rely on subjective reviews, limiting reproducibility and regulatory confidence. To address this gap, we developed an openly accessible DICOM dataset infused with synthetic PHI/PII and an evaluation framework for benchmarking image de-identification workflows. The Medical Image de-identification (MIDI) dataset was built using publicly available de-identified data from The Cancer Imaging Archive (TCIA). It includes 538 subjects (216 for validation, 322 for testing), 605 studies, 708 series, and 53,581 DICOM image instances. These span multiple vendors, imaging modalities, and cancer types. Synthetic PHI and PII were embedded into structured data elements, plain text data elements, and pixel data to simulate real-world identity leaks encountered by TCIA curation teams. Accompanying evaluation tools include a Python script, answer keys (known truth), and mapping files that enable automated comparison of curated data against expected transformations. The framework is aligned with the HIPAA Privacy Rule "Safe Harbor" method, DICOM PS3.15 Confidentiality Profiles, and TCIA best practices. It supports objective, standards-driven evaluation of de-identification workflows, promoting safer and more consistent medical image sharing.
title Medical Image De-Identification Resources: Synthetic DICOM Data and Tools for Validation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2508.01889