AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Xingjian, Wu, Qifeng, Ubaradka, Adithya S., Ding, Yiran, Que, Colleen, Jiang, Runmin, Xing, Jianhua, Wang, Tianyang, Xu, Min
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866914074366836736
author Li, Xingjian
Wu, Qifeng
Ubaradka, Adithya S.
Ding, Yiran
Que, Colleen
Jiang, Runmin
Xing, Jianhua
Wang, Tianyang
Xu, Min
author_facet Li, Xingjian
Wu, Qifeng
Ubaradka, Adithya S.
Ding, Yiran
Que, Colleen
Jiang, Runmin
Xing, Jianhua
Wang, Tianyang
Xu, Min
contents Medical image segmentation is vital for clinical diagnosis, yet current deep learning methods often demand extensive expert effort, i.e., either through annotating large training datasets or providing prompts at inference time for each new case. This paper introduces a zero-shot and automatic segmentation pipeline that combines off-the-shelf vision-language and segmentation foundation models. Given a medical image and a task definition (e.g., "segment the optic disc in an eye fundus image"), our method uses a grounding model to generate an initial bounding box, followed by a visual prompt boosting module that enhance the prompts, which are then processed by a promptable segmentation model to produce the final mask. To address the challenges of domain gap and result verification, we introduce a test-time adaptation framework featuring a set of learnable adaptors that align the medical inputs with foundation model representations. Its hyperparameters are optimized via Bayesian Optimization, guided by a proxy validation model without requiring ground-truth labels. Our pipeline offers an annotation-efficient and scalable solution for zero-shot medical image segmentation across diverse tasks. Our pipeline is evaluated on seven diverse medical imaging datasets and shows promising results. By proper decomposition and test-time adaptation, our fully automatic pipeline not only substantially surpasses the previously best-performing method, yielding a 69\% relative improvement in accuracy (Dice Score from 42.53 to 71.81), but also performs competitively with weakly-prompted interactive foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17931
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models
Li, Xingjian
Wu, Qifeng
Ubaradka, Adithya S.
Ding, Yiran
Que, Colleen
Jiang, Runmin
Xing, Jianhua
Wang, Tianyang
Xu, Min
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Medical image segmentation is vital for clinical diagnosis, yet current deep learning methods often demand extensive expert effort, i.e., either through annotating large training datasets or providing prompts at inference time for each new case. This paper introduces a zero-shot and automatic segmentation pipeline that combines off-the-shelf vision-language and segmentation foundation models. Given a medical image and a task definition (e.g., "segment the optic disc in an eye fundus image"), our method uses a grounding model to generate an initial bounding box, followed by a visual prompt boosting module that enhance the prompts, which are then processed by a promptable segmentation model to produce the final mask. To address the challenges of domain gap and result verification, we introduce a test-time adaptation framework featuring a set of learnable adaptors that align the medical inputs with foundation model representations. Its hyperparameters are optimized via Bayesian Optimization, guided by a proxy validation model without requiring ground-truth labels. Our pipeline offers an annotation-efficient and scalable solution for zero-shot medical image segmentation across diverse tasks. Our pipeline is evaluated on seven diverse medical imaging datasets and shows promising results. By proper decomposition and test-time adaptation, our fully automatic pipeline not only substantially surpasses the previously best-performing method, yielding a 69\% relative improvement in accuracy (Dice Score from 42.53 to 71.81), but also performs competitively with weakly-prompted interactive foundation models.
title AutoMiSeg: Automatic Medical Image Segmentation via Test-Time Adaptation of Foundation Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.17931