Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909010799624192 |
|---|---|
| author | Qu, Jingguo Han, Xinyang Ai, Jia Wu, Juan Zhao, Tong Xiao, Tonghuan Ning, Sheng Yang, Yuqi Qin, Jing King, Ann Dorothy Chu, Winnie Chiu-Wing Cai, Jing Ying, Michael Tin-Cheung |
| author_facet | Qu, Jingguo Han, Xinyang Ai, Jia Wu, Juan Zhao, Tong Xiao, Tonghuan Ning, Sheng Yang, Yuqi Qin, Jing King, Ann Dorothy Chu, Winnie Chiu-Wing Cai, Jing Ying, Michael Tin-Cheung |
| contents | Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by speckle noise, shadowing, and heterogeneous textures, often degrades the performance of off-the-shelf VLFMs. To bridge this gap, we propose a novel Hybrid Tuning (HT) strategy for the parameter-efficient adaptation of CLIP-based models to ultrasound analysis. Instead of updating the pre-trained weights, HT freezes the visual backbone and integrates a specialized lightweight adapter. This adapter features a Frequency Filtering module to suppress domain-specific periodic artifacts and a Noise Estimation module to dynamically calibrate feature representations. Extensive evaluations across six multi-center datasets demonstrate that our HT-enhanced models significantly outperform existing state-of-the-art adapters and medical VLFMs in both segmentation and classification tasks. Notably, HT exhibits exceptional data efficiency in few-shot scenarios and robust cross-dataset generalization. Our findings prove that preserving pre-trained semantic priors while explicitly modeling ultrasound-specific noise is key to unlocking foundational intelligence in automated ultrasound diagnosis. The source code is available at https://github.com/jinggqu/NextGen-UIA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_08849 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis Qu, Jingguo Han, Xinyang Ai, Jia Wu, Juan Zhao, Tong Xiao, Tonghuan Ning, Sheng Yang, Yuqi Qin, Jing King, Ann Dorothy Chu, Winnie Chiu-Wing Cai, Jing Ying, Michael Tin-Cheung Computer Vision and Pattern Recognition Vision-Language Foundation Models (VLFMs) exhibit remarkable generalization, yet their direct application to medical ultrasound is severely hindered by a profound modality gap. The unique acoustic physics of ultrasound, characterized by speckle noise, shadowing, and heterogeneous textures, often degrades the performance of off-the-shelf VLFMs. To bridge this gap, we propose a novel Hybrid Tuning (HT) strategy for the parameter-efficient adaptation of CLIP-based models to ultrasound analysis. Instead of updating the pre-trained weights, HT freezes the visual backbone and integrates a specialized lightweight adapter. This adapter features a Frequency Filtering module to suppress domain-specific periodic artifacts and a Noise Estimation module to dynamically calibrate feature representations. Extensive evaluations across six multi-center datasets demonstrate that our HT-enhanced models significantly outperform existing state-of-the-art adapters and medical VLFMs in both segmentation and classification tasks. Notably, HT exhibits exceptional data efficiency in few-shot scenarios and robust cross-dataset generalization. Our findings prove that preserving pre-trained semantic priors while explicitly modeling ultrasound-specific noise is key to unlocking foundational intelligence in automated ultrasound diagnosis. The source code is available at https://github.com/jinggqu/NextGen-UIA. |
| title | Adapting Vision-Language Foundation Model for Next Generation Medical Ultrasound Image Analysis |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2506.08849 |