Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yichi, Xiao, Feiyang, Xue, Le, Zhang, Wenbo, Feng, Gang, Zheng, Chenguang, Qi, Yuan, Cheng, Yuan, Hu, Zixin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915784222048256
author Zhang, Yichi
Xiao, Feiyang
Xue, Le
Zhang, Wenbo
Feng, Gang
Zheng, Chenguang
Qi, Yuan
Cheng, Yuan
Hu, Zixin
author_facet Zhang, Yichi
Xiao, Feiyang
Xue, Le
Zhang, Wenbo
Feng, Gang
Zheng, Chenguang
Qi, Yuan
Cheng, Yuan
Hu, Zixin
contents While emerging 3D medical foundation models are envisioned as versatile tools with offer general-purpose capabilities, their validation remains largely confined to regional and structural imaging, leaving a significant modality discrepancy unexplored. To provide a rigorous and objective assessment, we curate the UMD dataset comprising 490 whole-body PET/CT and 464 whole-body PET/MRI scans ($\sim$675k 2D images, $\sim$12k 3D organ annotations) and conduct a thorough and comprehensive evaluation of representative 3D segmentation foundation models. Through intra-subject controlled comparisons of paired scans, we isolate imaging modality as the primary independent variable to evaluate model robustness in real-world applications. Our evaluation reveals a stark discrepancy between literature-reported benchmarks and real-world efficacy, particularly when transitioning from structural to functional domains. Such systemic failures underscore that current 3D foundation models are far from achieving truly general-purpose status, necessitating a paradigm shift toward multi-modal training and evaluation to bridge the gap between idealized benchmarking and comprehensive clinical utility. This dataset and analysis establish a foundational cornerstone for future research to develop truly modality-agnostic medical foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07643
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation
Zhang, Yichi
Xiao, Feiyang
Xue, Le
Zhang, Wenbo
Feng, Gang
Zheng, Chenguang
Qi, Yuan
Cheng, Yuan
Hu, Zixin
Computer Vision and Pattern Recognition
While emerging 3D medical foundation models are envisioned as versatile tools with offer general-purpose capabilities, their validation remains largely confined to regional and structural imaging, leaving a significant modality discrepancy unexplored. To provide a rigorous and objective assessment, we curate the UMD dataset comprising 490 whole-body PET/CT and 464 whole-body PET/MRI scans ($\sim$675k 2D images, $\sim$12k 3D organ annotations) and conduct a thorough and comprehensive evaluation of representative 3D segmentation foundation models. Through intra-subject controlled comparisons of paired scans, we isolate imaging modality as the primary independent variable to evaluate model robustness in real-world applications. Our evaluation reveals a stark discrepancy between literature-reported benchmarks and real-world efficacy, particularly when transitioning from structural to functional domains. Such systemic failures underscore that current 3D foundation models are far from achieving truly general-purpose status, necessitating a paradigm shift toward multi-modal training and evaluation to bridge the gap between idealized benchmarking and comprehensive clinical utility. This dataset and analysis establish a foundational cornerstone for future research to develop truly modality-agnostic medical foundation models.
title Uncovering Modality Discrepancy and Generalization Illusion for General-Purpose 3D Medical Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.07643