In search of truth: Evaluating concordance of AI-based anatomy segmentation models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Giebeler, Lena, Krishnaswamy, Deepa, Clunie, David, Wasserthal, Jakob, Sundar, Lalith Kumar Shiyam, Diaz-Pinto, Andres, Maier-Hein, Klaus H., Xu, Murong, Menze, Bjoern, Pieper, Steve, Kikinis, Ron, Fedorov, Andrey
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866915918462844928
author Giebeler, Lena
Krishnaswamy, Deepa
Clunie, David
Wasserthal, Jakob
Sundar, Lalith Kumar Shiyam
Diaz-Pinto, Andres
Maier-Hein, Klaus H.
Xu, Murong
Menze, Bjoern
Pieper, Steve
Kikinis, Ron
Fedorov, Andrey
author_facet Giebeler, Lena
Krishnaswamy, Deepa
Clunie, David
Wasserthal, Jakob
Sundar, Lalith Kumar Shiyam
Diaz-Pinto, Andres
Maier-Hein, Klaus H.
Xu, Murong
Menze, Bjoern
Pieper, Steve
Kikinis, Ron
Fedorov, Andrey
contents Purpose AI-based methods for anatomy segmentation can help automate characterization of large imaging datasets. The growing number of similar in functionality models raises the challenge of evaluating them on datasets that do not contain ground truth annotations. We introduce a practical framework to assist in this task. Approach We harmonize the segmentation results into a standard, interoperable representation, which enables consistent, terminology-based labeling of the structures. We extend 3D Slicer to streamline loading and comparison of these harmonized segmentations, and demonstrate how standard representation simplifies review of the results using interactive summary plots and browser-based visualization using OHIF Viewer. To demonstrate the utility of the approach we apply it to evaluating segmentation of 31 anatomical structures (lungs, vertebrae, ribs, and heart) by six open-source models - TotalSegmentator 1.5 and 2.6, Auto3DSeg, MOOSE, MultiTalent, and CADS - for a sample of Computed Tomography (CT) scans from the publicly available National Lung Screening Trial (NLST) dataset. Results We demonstrate the utility of the framework in enabling automating loading, structure-wise inspection and comparison across models. Preliminary results ascertain practical utility of the approach in allowing quick detection and review of problematic results. The comparison shows excellent agreement segmenting some (e.g., lung) but not all structures (e.g., some models produce invalid vertebrae or rib segmentations). Conclusions The resources developed are linked from https://imagingdatacommons.github.io/segmentation-comparison/ including segmentation harmonization scripts, summary plots, and visualization tools. This work assists in model evaluation in absence of ground truth, ultimately enabling informed model selection.
format Preprint
id arxiv_https___arxiv_org_abs_2512_15921
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle In search of truth: Evaluating concordance of AI-based anatomy segmentation models
Giebeler, Lena
Krishnaswamy, Deepa
Clunie, David
Wasserthal, Jakob
Sundar, Lalith Kumar Shiyam
Diaz-Pinto, Andres
Maier-Hein, Klaus H.
Xu, Murong
Menze, Bjoern
Pieper, Steve
Kikinis, Ron
Fedorov, Andrey
Image and Video Processing
Computer Vision and Pattern Recognition
Purpose AI-based methods for anatomy segmentation can help automate characterization of large imaging datasets. The growing number of similar in functionality models raises the challenge of evaluating them on datasets that do not contain ground truth annotations. We introduce a practical framework to assist in this task. Approach We harmonize the segmentation results into a standard, interoperable representation, which enables consistent, terminology-based labeling of the structures. We extend 3D Slicer to streamline loading and comparison of these harmonized segmentations, and demonstrate how standard representation simplifies review of the results using interactive summary plots and browser-based visualization using OHIF Viewer. To demonstrate the utility of the approach we apply it to evaluating segmentation of 31 anatomical structures (lungs, vertebrae, ribs, and heart) by six open-source models - TotalSegmentator 1.5 and 2.6, Auto3DSeg, MOOSE, MultiTalent, and CADS - for a sample of Computed Tomography (CT) scans from the publicly available National Lung Screening Trial (NLST) dataset. Results We demonstrate the utility of the framework in enabling automating loading, structure-wise inspection and comparison across models. Preliminary results ascertain practical utility of the approach in allowing quick detection and review of problematic results. The comparison shows excellent agreement segmenting some (e.g., lung) but not all structures (e.g., some models produce invalid vertebrae or rib segmentations). Conclusions The resources developed are linked from https://imagingdatacommons.github.io/segmentation-comparison/ including segmentation harmonization scripts, summary plots, and visualization tools. This work assists in model evaluation in absence of ground truth, ultimately enabling informed model selection.
title In search of truth: Evaluating concordance of AI-based anatomy segmentation models
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.15921