Saved in:
Bibliographic Details
Main Authors: Weng, Taohan, Hu, Kaibing, Liu, Henan, Liu, Siya, Liu, Xiaoyang, Liu, Zhenyu, Ren, Jiren, Wang, Boyan, Wang, Boyang, Wang, Yiyu, Wu, Yalun, Yan, Chaoran, Yan, Kaiwen, Yu, Jinze, Zhang, Chi, Zhang, Duo, Zheng, Haoyun, Guo, Xiaoqing, Souquet, Jacques, Guo, Hongcheng, Le, Anjie
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2509.25748
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912659167772672
author Weng, Taohan
Hu, Kaibing
Liu, Henan
Liu, Siya
Liu, Xiaoyang
Liu, Zhenyu
Ren, Jiren
Wang, Boyan
Wang, Boyang
Wang, Yiyu
Wu, Yalun
Yan, Chaoran
Yan, Kaiwen
Yu, Jinze
Zhang, Chi
Zhang, Duo
Zheng, Haoyun
Guo, Xiaoqing
Souquet, Jacques
Guo, Hongcheng
Le, Anjie
author_facet Weng, Taohan
Hu, Kaibing
Liu, Henan
Liu, Siya
Liu, Xiaoyang
Liu, Zhenyu
Ren, Jiren
Wang, Boyan
Wang, Boyang
Wang, Yiyu
Wu, Yalun
Yan, Chaoran
Yan, Kaiwen
Yu, Jinze
Zhang, Chi
Zhang, Duo
Zheng, Haoyun
Guo, Xiaoqing
Souquet, Jacques
Guo, Hongcheng
Le, Anjie
contents Ultrasound is crucial in modern medicine but faces challenges like operator dependence, image noise, and real-time scanning, hindering AI integration. While large multimodal models excel in other medical imaging areas, they struggle with ultrasound's complexities. To address this, we introduce Dolphin v1.0 (V1) and its reasoning-augmented version, Dolphin R1-the first large-scale multimodal ultrasound foundation models unifying diverse clinical tasks in a single vision-language framework.To tackle ultrasound variability and noise, we curated a 2-million-scale multimodal dataset, combining textbook knowledge, public data, synthetic samples, and general corpora. This ensures robust perception, generalization, and clinical adaptability.The Dolphin series employs a three-stage training strategy: domain-specialized pretraining, instruction-driven alignment, and reinforcement-based refinement. Dolphin v1.0 delivers reliable performance in classification, detection, regression, and report generation. Dolphin R1 enhances diagnostic inference, reasoning transparency, and interpretability through reinforcement learning with ultrasound-specific rewards.Evaluated on U2-Bench across eight ultrasound tasks, Dolphin R1 achieves a U2-score of 0.5835-over twice the second-best model (0.2968) setting a new state of the art. Dolphin v1.0 also performs competitively, validating the unified framework. Comparisons show reasoning-enhanced training significantly improves diagnostic accuracy, consistency, and interpretability, highlighting its importance for high-stakes medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25748
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dolphin v1.0 Technical Report
Weng, Taohan
Hu, Kaibing
Liu, Henan
Liu, Siya
Liu, Xiaoyang
Liu, Zhenyu
Ren, Jiren
Wang, Boyan
Wang, Boyang
Wang, Yiyu
Wu, Yalun
Yan, Chaoran
Yan, Kaiwen
Yu, Jinze
Zhang, Chi
Zhang, Duo
Zheng, Haoyun
Guo, Xiaoqing
Souquet, Jacques
Guo, Hongcheng
Le, Anjie
Computer Vision and Pattern Recognition
Artificial Intelligence
Ultrasound is crucial in modern medicine but faces challenges like operator dependence, image noise, and real-time scanning, hindering AI integration. While large multimodal models excel in other medical imaging areas, they struggle with ultrasound's complexities. To address this, we introduce Dolphin v1.0 (V1) and its reasoning-augmented version, Dolphin R1-the first large-scale multimodal ultrasound foundation models unifying diverse clinical tasks in a single vision-language framework.To tackle ultrasound variability and noise, we curated a 2-million-scale multimodal dataset, combining textbook knowledge, public data, synthetic samples, and general corpora. This ensures robust perception, generalization, and clinical adaptability.The Dolphin series employs a three-stage training strategy: domain-specialized pretraining, instruction-driven alignment, and reinforcement-based refinement. Dolphin v1.0 delivers reliable performance in classification, detection, regression, and report generation. Dolphin R1 enhances diagnostic inference, reasoning transparency, and interpretability through reinforcement learning with ultrasound-specific rewards.Evaluated on U2-Bench across eight ultrasound tasks, Dolphin R1 achieves a U2-score of 0.5835-over twice the second-best model (0.2968) setting a new state of the art. Dolphin v1.0 also performs competitively, validating the unified framework. Comparisons show reasoning-enhanced training significantly improves diagnostic accuracy, consistency, and interpretability, highlighting its importance for high-stakes medical AI.
title Dolphin v1.0 Technical Report
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2509.25748