Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915784222048256 |
|---|---|
| author | Zhang, Yichi Xiao, Feiyang Xue, Le Zhang, Wenbo Feng, Gang Zheng, Chenguang Qi, Yuan Cheng, Yuan Hu, Zixin |
| author_facet | Zhang, Yichi Xiao, Feiyang Xue, Le Zhang, Wenbo Feng, Gang Zheng, Chenguang Qi, Yuan Cheng, Yuan Hu, Zixin |
| contents | While emerging 3D medical foundation models are envisioned as versatile tools with offer general-purpose capabilities, their validation remains largely confined to regional and structural imaging, leaving a significant modality discrepancy unexplored. To provide a rigorous and objective assessment, we curate the UMD dataset comprising 490 whole-body PET/CT and 464 whole-body PET/MRI scans ($\sim$675k 2D images, $\sim$12k 3D organ annotations) and conduct a thorough and comprehensive evaluation of representative 3D segmentation foundation models. Through intra-subject controlled comparisons of paired scans, we isolate imaging modality as the primary independent variable to evaluate model robustness in real-world applications. Our evaluation reveals a stark discrepancy between literature-reported benchmarks and real-world efficacy, particularly when transitioning from structural to functional domains. Such systemic failures underscore that current 3D foundation models are far from achieving truly general-purpose status, necessitating a paradigm shift toward multi-modal training and evaluation to bridge the gap between idealized benchmarking and comprehensive clinical utility. This dataset and analysis establish a foundational cornerstone for future research to develop truly modality-agnostic medical foundation models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2602_07643 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation Zhang, Yichi Xiao, Feiyang Xue, Le Zhang, Wenbo Feng, Gang Zheng, Chenguang Qi, Yuan Cheng, Yuan Hu, Zixin Computer Vision and Pattern Recognition While emerging 3D medical foundation models are envisioned as versatile tools with offer general-purpose capabilities, their validation remains largely confined to regional and structural imaging, leaving a significant modality discrepancy unexplored. To provide a rigorous and objective assessment, we curate the UMD dataset comprising 490 whole-body PET/CT and 464 whole-body PET/MRI scans ($\sim$675k 2D images, $\sim$12k 3D organ annotations) and conduct a thorough and comprehensive evaluation of representative 3D segmentation foundation models. Through intra-subject controlled comparisons of paired scans, we isolate imaging modality as the primary independent variable to evaluate model robustness in real-world applications. Our evaluation reveals a stark discrepancy between literature-reported benchmarks and real-world efficacy, particularly when transitioning from structural to functional domains. Such systemic failures underscore that current 3D foundation models are far from achieving truly general-purpose status, necessitating a paradigm shift toward multi-modal training and evaluation to bridge the gap between idealized benchmarking and comprehensive clinical utility. This dataset and analysis establish a foundational cornerstone for future research to develop truly modality-agnostic medical foundation models. |
| title | Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2602.07643 |