_version_ 1866917961598500864
author Garrucho, Lidia
Kushibar, Kaisar
Reidel, Claire-Anne
Joshi, Smriti
Osuala, Richard
Tsirikoglou, Apostolia
Bobowicz, Maciej
del Riego, Javier
Catanese, Alessandro
Gwoździewicz, Katarzyna
Cosaka, Maria-Laura
Abo-Elhoda, Pasant M.
Tantawy, Sara W.
Sakrana, Shorouq S.
Shawky-Abdelfatah, Norhan O.
Abdo-Salem, Amr Muhammad
Kozana, Androniki
Divjak, Eugen
Ivanac, Gordana
Nikiforaki, Katerina
Klontzas, Michail E.
García-Dosdá, Rosa
Gulsun-Akpinar, Meltem
Lafcı, Oğuz
Mann, Ritse
Martín-Isla, Carlos
Prior, Fred
Marias, Kostas
Starmans, Martijn P. A.
Strand, Fredrik
Díaz, Oliver
Igual, Laura
Lekadir, Karim
author_facet Garrucho, Lidia
Kushibar, Kaisar
Reidel, Claire-Anne
Joshi, Smriti
Osuala, Richard
Tsirikoglou, Apostolia
Bobowicz, Maciej
del Riego, Javier
Catanese, Alessandro
Gwoździewicz, Katarzyna
Cosaka, Maria-Laura
Abo-Elhoda, Pasant M.
Tantawy, Sara W.
Sakrana, Shorouq S.
Shawky-Abdelfatah, Norhan O.
Abdo-Salem, Amr Muhammad
Kozana, Androniki
Divjak, Eugen
Ivanac, Gordana
Nikiforaki, Katerina
Klontzas, Michail E.
García-Dosdá, Rosa
Gulsun-Akpinar, Meltem
Lafcı, Oğuz
Mann, Ritse
Martín-Isla, Carlos
Prior, Fred
Marias, Kostas
Starmans, Martijn P. A.
Strand, Fredrik
Díaz, Oliver
Igual, Laura
Lekadir, Karim
contents Artificial Intelligence (AI) research in breast cancer Magnetic Resonance Imaging (MRI) faces challenges due to limited expert-labeled segmentations. To address this, we present a multicenter dataset of 1506 pre-treatment T1-weighted dynamic contrast-enhanced MRI cases, including expert annotations of primary tumors and non-mass-enhanced regions. The dataset integrates imaging data from four collections in The Cancer Imaging Archive (TCIA), where only 163 cases with expert segmentations were initially available. To facilitate the annotation process, a deep learning model was trained to produce preliminary segmentations for the remaining cases. These were subsequently corrected and verified by 16 breast cancer experts (averaging 9 years of experience), creating a fully annotated dataset. Additionally, the dataset includes 49 harmonized clinical and demographic variables, as well as pre-trained weights for a baseline nnU-Net model trained on the annotated data. This resource addresses a critical gap in publicly available breast cancer datasets, enabling the development, validation, and benchmarking of advanced deep learning models, thus driving progress in breast cancer diagnostics, treatment response prediction, and personalized care.
format Preprint
id arxiv_https___arxiv_org_abs_2406_13844
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A large-scale multicenter breast cancer DCE-MRI benchmark dataset with expert segmentations
Garrucho, Lidia
Kushibar, Kaisar
Reidel, Claire-Anne
Joshi, Smriti
Osuala, Richard
Tsirikoglou, Apostolia
Bobowicz, Maciej
del Riego, Javier
Catanese, Alessandro
Gwoździewicz, Katarzyna
Cosaka, Maria-Laura
Abo-Elhoda, Pasant M.
Tantawy, Sara W.
Sakrana, Shorouq S.
Shawky-Abdelfatah, Norhan O.
Abdo-Salem, Amr Muhammad
Kozana, Androniki
Divjak, Eugen
Ivanac, Gordana
Nikiforaki, Katerina
Klontzas, Michail E.
García-Dosdá, Rosa
Gulsun-Akpinar, Meltem
Lafcı, Oğuz
Mann, Ritse
Martín-Isla, Carlos
Prior, Fred
Marias, Kostas
Starmans, Martijn P. A.
Strand, Fredrik
Díaz, Oliver
Igual, Laura
Lekadir, Karim
Computer Vision and Pattern Recognition
Artificial Intelligence
Databases
Artificial Intelligence (AI) research in breast cancer Magnetic Resonance Imaging (MRI) faces challenges due to limited expert-labeled segmentations. To address this, we present a multicenter dataset of 1506 pre-treatment T1-weighted dynamic contrast-enhanced MRI cases, including expert annotations of primary tumors and non-mass-enhanced regions. The dataset integrates imaging data from four collections in The Cancer Imaging Archive (TCIA), where only 163 cases with expert segmentations were initially available. To facilitate the annotation process, a deep learning model was trained to produce preliminary segmentations for the remaining cases. These were subsequently corrected and verified by 16 breast cancer experts (averaging 9 years of experience), creating a fully annotated dataset. Additionally, the dataset includes 49 harmonized clinical and demographic variables, as well as pre-trained weights for a baseline nnU-Net model trained on the annotated data. This resource addresses a critical gap in publicly available breast cancer datasets, enabling the development, validation, and benchmarking of advanced deep learning models, thus driving progress in breast cancer diagnostics, treatment response prediction, and personalized care.
title A large-scale multicenter breast cancer DCE-MRI benchmark dataset with expert segmentations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Databases
url https://arxiv.org/abs/2406.13844