Refining 3D Medical Segmentation with Verbal Instruction

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Xie, Kangxian, Yang, Jiancheng, Pinter, Nandor, Wu, Chao, Bozorgtabar, Behzad, Gao, Mingchen
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915864661458944
author Xie, Kangxian
Yang, Jiancheng
Pinter, Nandor
Wu, Chao
Bozorgtabar, Behzad
Gao, Mingchen
author_facet Xie, Kangxian
Yang, Jiancheng
Pinter, Nandor
Wu, Chao
Bozorgtabar, Behzad
Gao, Mingchen
contents Accurate 3D anatomical segmentation is essential for clinical diagnosis and surgical planning. However, automated models frequently generate suboptimal shape predictions due to factors such as limited and imbalanced training data, inadequate labeling quality, and distribution shifts between training and deployment settings. A natural solution is to iteratively refine the predicted shape based on the radiologists' verbal instructions. However, this is hindered by the scarcity of paired data that explicitly links erroneous shapes to corresponding corrective instructions. As an initial step toward addressing this limitation, we introduce CoWTalk, a benchmark comprising 3D arterial anatomies with controllable synthesized anatomical errors and their corresponding repairing instructions. Building on this benchmark, we further propose an iterative refinement model that represents 3D shapes as vector sets and interacts with textual instructions to progressively update the target shape. Experimental results demonstrate that our method achieves significant improvements over corrupted inputs and competitive baselines, highlighting the feasibility of language-driven clinician-in-the-loop refinement for 3D medical shapes modeling.
format Preprint
id arxiv_https___arxiv_org_abs_2603_14496
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Refining 3D Medical Segmentation with Verbal Instruction
Xie, Kangxian
Yang, Jiancheng
Pinter, Nandor
Wu, Chao
Bozorgtabar, Behzad
Gao, Mingchen
Computer Vision and Pattern Recognition
Machine Learning
Accurate 3D anatomical segmentation is essential for clinical diagnosis and surgical planning. However, automated models frequently generate suboptimal shape predictions due to factors such as limited and imbalanced training data, inadequate labeling quality, and distribution shifts between training and deployment settings. A natural solution is to iteratively refine the predicted shape based on the radiologists' verbal instructions. However, this is hindered by the scarcity of paired data that explicitly links erroneous shapes to corresponding corrective instructions. As an initial step toward addressing this limitation, we introduce CoWTalk, a benchmark comprising 3D arterial anatomies with controllable synthesized anatomical errors and their corresponding repairing instructions. Building on this benchmark, we further propose an iterative refinement model that represents 3D shapes as vector sets and interacts with textual instructions to progressively update the target shape. Experimental results demonstrate that our method achieves significant improvements over corrupted inputs and competitive baselines, highlighting the feasibility of language-driven clinician-in-the-loop refinement for 3D medical shapes modeling.
title Refining 3D Medical Segmentation with Verbal Instruction
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2603.14496