Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xiaoran, Chen, Eric Z., Zhao, Lin, Chen, Xiao, Liu, Yikang, Maihe, Boris, Duncan, James S., Chen, Terrence, Sun, Shanhui
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909559879106560
author Zhang, Xiaoran
Chen, Eric Z.
Zhao, Lin
Chen, Xiao
Liu, Yikang
Maihe, Boris
Duncan, James S.
Chen, Terrence
Sun, Shanhui
author_facet Zhang, Xiaoran
Chen, Eric Z.
Zhao, Lin
Chen, Xiao
Liu, Yikang
Maihe, Boris
Duncan, James S.
Chen, Terrence
Sun, Shanhui
contents We propose a novel approach that adapts hierarchical vision foundation models for real-time ultrasound image segmentation. Existing ultrasound segmentation methods often struggle with adaptability to new tasks, relying on costly manual annotations, while real-time approaches generally fail to match state-of-the-art performance. To overcome these limitations, we introduce an adaptive framework that leverages the vision foundation model Hiera to extract multi-scale features, interleaved with DINOv2 representations to enhance visual expressiveness. These enriched features are then decoded to produce precise and robust segmentation. We conduct extensive evaluations on six public datasets and one in-house dataset, covering both cardiac and thyroid ultrasound segmentation. Experiments show that our approach outperforms state-of-the-art methods across multiple datasets and excels with limited supervision, surpassing nnUNet by over 20\% on average in the 1\% and 10\% data settings. Our method achieves $\sim$77 FPS inference speed with TensorRT on a single GPU, enabling real-time clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2503_24368
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
Zhang, Xiaoran
Chen, Eric Z.
Zhao, Lin
Chen, Xiao
Liu, Yikang
Maihe, Boris
Duncan, James S.
Chen, Terrence
Sun, Shanhui
Computer Vision and Pattern Recognition
We propose a novel approach that adapts hierarchical vision foundation models for real-time ultrasound image segmentation. Existing ultrasound segmentation methods often struggle with adaptability to new tasks, relying on costly manual annotations, while real-time approaches generally fail to match state-of-the-art performance. To overcome these limitations, we introduce an adaptive framework that leverages the vision foundation model Hiera to extract multi-scale features, interleaved with DINOv2 representations to enhance visual expressiveness. These enriched features are then decoded to produce precise and robust segmentation. We conduct extensive evaluations on six public datasets and one in-house dataset, covering both cardiac and thyroid ultrasound segmentation. Experiments show that our approach outperforms state-of-the-art methods across multiple datasets and excels with limited supervision, surpassing nnUNet by over 20\% on average in the 1\% and 10\% data settings. Our method achieves $\sim$77 FPS inference speed with TensorRT on a single GPU, enabling real-time clinical applications.
title Adapting Vision Foundation Models for Real-time Ultrasound Image Segmentation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2503.24368