Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ndir, Tidiane Camaret, Pfefferle, Alexander, Schirrmeister, Robin Tibor
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917148103802880
author Ndir, Tidiane Camaret
Pfefferle, Alexander
Schirrmeister, Robin Tibor
author_facet Ndir, Tidiane Camaret
Pfefferle, Alexander
Schirrmeister, Robin Tibor
contents Interactive 3D biomedical image segmentation requires efficient models that can iteratively refine predictions based on user prompts. Current foundation models either lack volumetric awareness or suffer from limited interactive capabilities. We propose a training strategy that combines dynamic volumetric prompt generation with content-aware adaptive cropping to optimize the use of the image encoder. Our method simulates realistic user interaction patterns during training while addressing the computational challenges of learning from sequential refinement feedback on a single GPU. For efficient training, we initialize our network using the publicly available weights from the nnInteractive segmentation model. Evaluation on the \textbf{Foundation Models for Interactive 3D Biomedical Image Segmentation} competition demonstrates strong performance with an average final Dice score of 0.6385, normalized surface distance of 0.6614, and area-under-the-curve metrics of 2.4799 (Dice) and 2.5671 (NSD).
format Preprint
id arxiv_https___arxiv_org_abs_2510_03189
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training
Ndir, Tidiane Camaret
Pfefferle, Alexander
Schirrmeister, Robin Tibor
Computer Vision and Pattern Recognition
Interactive 3D biomedical image segmentation requires efficient models that can iteratively refine predictions based on user prompts. Current foundation models either lack volumetric awareness or suffer from limited interactive capabilities. We propose a training strategy that combines dynamic volumetric prompt generation with content-aware adaptive cropping to optimize the use of the image encoder. Our method simulates realistic user interaction patterns during training while addressing the computational challenges of learning from sequential refinement feedback on a single GPU. For efficient training, we initialize our network using the publicly available weights from the nnInteractive segmentation model. Evaluation on the \textbf{Foundation Models for Interactive 3D Biomedical Image Segmentation} competition demonstrates strong performance with an average final Dice score of 0.6385, normalized surface distance of 0.6614, and area-under-the-curve metrics of 2.4799 (Dice) and 2.5671 (NSD).
title Dynamic Prompt Generation for Interactive 3D Medical Image Segmentation Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.03189