Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zeng, Chengxi, Smithard, David, Gambaruto, Alberto M, Burghardt, Tilo
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912211243368448
author Zeng, Chengxi
Smithard, David
Gambaruto, Alberto M
Burghardt, Tilo
author_facet Zeng, Chengxi
Smithard, David
Gambaruto, Alberto M
Burghardt, Tilo
contents Vision foundation models have demonstrated exceptional generalization capabilities in segmentation tasks for both generic and specialized images. However, a performance gap persists between foundation models and task-specific, specialized models. Fine-tuning foundation models on downstream datasets is often necessary to bridge this gap. Unfortunately, obtaining fully annotated ground truth for downstream datasets is both challenging and costly. To address this limitation, we propose a novel test-time training paradigm that enhances the performance of foundation models on downstream datasets without requiring full annotations. Specifically, our method employs simple point prompts to guide a test-time semi-self-supervised training task. The model learns by resolving the ambiguity of the point prompt through various augmentations. This approach directly tackles challenges in the medical imaging field, where acquiring annotations is both time-intensive and expensive. We conducted extensive experiments on our new Videofluoroscopy dataset (VFSS-5k) for the instance segmentation task, achieving an average Dice coefficient of 0.868 across 12 anatomies with a single model.
format Preprint
id arxiv_https___arxiv_org_abs_2501_18474
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations
Zeng, Chengxi
Smithard, David
Gambaruto, Alberto M
Burghardt, Tilo
Computer Vision and Pattern Recognition
Vision foundation models have demonstrated exceptional generalization capabilities in segmentation tasks for both generic and specialized images. However, a performance gap persists between foundation models and task-specific, specialized models. Fine-tuning foundation models on downstream datasets is often necessary to bridge this gap. Unfortunately, obtaining fully annotated ground truth for downstream datasets is both challenging and costly. To address this limitation, we propose a novel test-time training paradigm that enhances the performance of foundation models on downstream datasets without requiring full annotations. Specifically, our method employs simple point prompts to guide a test-time semi-self-supervised training task. The model learns by resolving the ambiguity of the point prompt through various augmentations. This approach directly tackles challenges in the medical imaging field, where acquiring annotations is both time-intensive and expensive. We conducted extensive experiments on our new Videofluoroscopy dataset (VFSS-5k) for the instance segmentation task, achieving an average Dice coefficient of 0.868 across 12 anatomies with a single model.
title Tuning Vision Foundation Model via Test-Time Prompt-Guided Training for VFSS Segmentations
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.18474