Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Jingguo, Han, Xinyang, Ai, Jia, Wu, Juan, Zhao, Tong, Xiao, Tonghuan, Ning, Sheng, Yang, Yuqi, Qin, Jing, King, Ann Dorothy, Chu, Winnie Chiu-Wing, Cai, Jing, Ying, Michael Tin-Cheung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909010799624192
author Qu, Jingguo
Han, Xinyang
Ai, Jia
Wu, Juan
Zhao, Tong
Xiao, Tonghuan
Ning, Sheng
Yang, Yuqi
Qin, Jing
King, Ann Dorothy
Chu, Winnie Chiu-Wing
Cai, Jing
Ying, Michael Tin-Cheung
author_facet Qu, Jingguo
Han, Xinyang
Ai, Jia
Wu, Juan
Zhao, Tong
Xiao, Tonghuan
Ning, Sheng
Yang, Yuqi
Qin, Jing
King, Ann Dorothy
Chu, Winnie Chiu-Wing
Cai, Jing
Ying, Michael Tin-Cheung
contents Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by speckle noise, shadowing, and heterogeneous textures, often degrades the performance of off-the-shelf VLFMs. To bridge this gap, we propose a novel Hybrid Tuning (HT) strategy for the parameter-efficient adaptation of CLIP-based models to ultrasound analysis. Instead of updating the pre-trained weights, HT freezes the visual backbone and integrates a specialized lightweight adapter. This adapter features a Frequency Filtering module to suppress domain-specific periodic artifacts and a Noise Estimation module to dynamically calibrate feature representations. Extensive evaluations across six multi-center datasets demonstrate that our HT-enhanced models significantly outperform existing state-of-the-art adapters and medical VLFMs in both segmentation and classification tasks. Notably, HT exhibits exceptional data efficiency in few-shot scenarios and robust cross-dataset generalization. Our findings prove that preserving pre-trained semantic priors while explicitly modeling ultrasound-specific noise is key to unlocking foundational intelligence in automated ultrasound diagnosis. The source code is available at https://github.com/jinggqu/NextGen-UIA.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08849
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
Qu, Jingguo
Han, Xinyang
Ai, Jia
Wu, Juan
Zhao, Tong
Xiao, Tonghuan
Ning, Sheng
Yang, Yuqi
Qin, Jing
King, Ann Dorothy
Chu, Winnie Chiu-Wing
Cai, Jing
Ying, Michael Tin-Cheung
Computer Vision and Pattern Recognition
Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by speckle noise, shadowing, and heterogeneous textures, often degrades the performance of off-the-shelf VLFMs. To bridge this gap, we propose a novel Hybrid Tuning (HT) strategy for the parameter-efficient adaptation of CLIP-based models to ultrasound analysis. Instead of updating the pre-trained weights, HT freezes the visual backbone and integrates a specialized lightweight adapter. This adapter features a Frequency Filtering module to suppress domain-specific periodic artifacts and a Noise Estimation module to dynamically calibrate feature representations. Extensive evaluations across six multi-center datasets demonstrate that our HT-enhanced models significantly outperform existing state-of-the-art adapters and medical VLFMs in both segmentation and classification tasks. Notably, HT exhibits exceptional data efficiency in few-shot scenarios and robust cross-dataset generalization. Our findings prove that preserving pre-trained semantic priors while explicitly modeling ultrasound-specific noise is key to unlocking foundational intelligence in automated ultrasound diagnosis. The source code is available at https://github.com/jinggqu/NextGen-UIA.
title Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.08849