V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Wang, Yuan, Liu, Jiaxiang, Gao, Shujian, Feng, Bin, Tang, Zhihang, Gai, Xiaotang, Wu, Jian, Liu, Zuozhu
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912452778655744
author Wang, Yuan
Liu, Jiaxiang
Gao, Shujian
Feng, Bin
Tang, Zhihang
Gai, Xiaotang
Wu, Jian
Liu, Zuozhu
author_facet Wang, Yuan
Liu, Jiaxiang
Gao, Shujian
Feng, Bin
Tang, Zhihang
Gai, Xiaotang
Wu, Jian
Liu, Zuozhu
contents Recent advances in multimodal techniques have led to significant progress in Medical Visual Question Answering (Med-VQA). However, most existing models focus on global image features rather than localizing disease-specific regions crucial for diagnosis. Additionally, current research tends to emphasize answer accuracy at the expense of the reasoning pathway, yet both are crucial for clinical decision-making. To address these challenges, we propose From Vision to Text Chain-of-Thought (V2T-CoT), a novel approach that automates the localization of preference areas within biomedical images and incorporates this localization into region-level pixel attention as knowledge for Vision CoT. By fine-tuning the vision language model on constructed R-Med 39K dataset, V2T-CoT provides definitive medical reasoning paths. V2T-CoT integrates visual grounding with textual rationale generation to establish precise and explainable diagnostic results. Experimental results across four Med-VQA benchmarks demonstrate state-of-the-art performance, achieving substantial improvements in both performance and interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_19610
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
Wang, Yuan
Liu, Jiaxiang
Gao, Shujian
Feng, Bin
Tang, Zhihang
Gai, Xiaotang
Wu, Jian
Liu, Zuozhu
Computational Engineering, Finance, and Science
Recent advances in multimodal techniques have led to significant progress in Medical Visual Question Answering (Med-VQA). However, most existing models focus on global image features rather than localizing disease-specific regions crucial for diagnosis. Additionally, current research tends to emphasize answer accuracy at the expense of the reasoning pathway, yet both are crucial for clinical decision-making. To address these challenges, we propose From Vision to Text Chain-of-Thought (V2T-CoT), a novel approach that automates the localization of preference areas within biomedical images and incorporates this localization into region-level pixel attention as knowledge for Vision CoT. By fine-tuning the vision language model on constructed R-Med 39K dataset, V2T-CoT provides definitive medical reasoning paths. V2T-CoT integrates visual grounding with textual rationale generation to establish precise and explainable diagnostic results. Experimental results across four Med-VQA benchmarks demonstrate state-of-the-art performance, achieving substantial improvements in both performance and interpretability.
title V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
topic Computational Engineering, Finance, and Science
url https://arxiv.org/abs/2506.19610