S-Chain: Structured Visual Chain-of-Thought For Medicine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Le-Duc, Khai, Nguyen, Duy M. H., Trinh, Phuong T. H., Nguyen, Tien-Phat, Diep, Nghiem T., Ngo, An, Vu, Tung, Vuong, Trinh, Nguyen, Anh-Tien, Nguyen, Mau, Hoang, Van Trung, Nguyen, Khai-Nguyen, Nguyen, Hy, Ngo, Chris, Liu, Anji, Ho, Nhat, Hauschild, Anne-Christin, Nguyen, Khanh Xuan, Nguyen-Tang, Thanh, Xie, Pengtao, Sonntag, Daniel, Zou, James, Niepert, Mathias, Nguyen, Anh Totti
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911233961099264
author Le-Duc, Khai
Nguyen, Duy M. H.
Trinh, Phuong T. H.
Nguyen, Tien-Phat
Diep, Nghiem T.
Ngo, An
Vu, Tung
Vuong, Trinh
Nguyen, Anh-Tien
Nguyen, Mau
Hoang, Van Trung
Nguyen, Khai-Nguyen
Nguyen, Hy
Ngo, Chris
Liu, Anji
Ho, Nhat
Hauschild, Anne-Christin
Nguyen, Khanh Xuan
Nguyen-Tang, Thanh
Xie, Pengtao
Sonntag, Daniel
Zou, James
Niepert, Mathias
Nguyen, Anh Totti
author_facet Le-Duc, Khai
Nguyen, Duy M. H.
Trinh, Phuong T. H.
Nguyen, Tien-Phat
Diep, Nghiem T.
Ngo, An
Vu, Tung
Vuong, Trinh
Nguyen, Anh-Tien
Nguyen, Mau
Hoang, Van Trung
Nguyen, Khai-Nguyen
Nguyen, Hy
Ngo, Chris
Liu, Anji
Ho, Nhat
Hauschild, Anne-Christin
Nguyen, Khanh Xuan
Nguyen-Tang, Thanh
Xie, Pengtao
Sonntag, Daniel
Zou, James
Niepert, Mathias
Nguyen, Anh Totti
contents Faithful reasoning in medical vision-language models (VLMs) requires not only accurate predictions but also transparent alignment between textual rationales and visual evidence. While Chain-of-Thought (CoT) prompting has shown promise in medical visual question answering (VQA), no large-scale expert-level dataset has captured stepwise reasoning with precise visual grounding. We introduce S-Chain, the first large-scale dataset of 12,000 expert-annotated medical images with bounding boxes and structured visual CoT (SV-CoT), explicitly linking visual regions to reasoning steps. The dataset further supports 16 languages, totaling over 700k VQA pairs for broad multilingual applicability. Using S-Chain, we benchmark state-of-the-art medical VLMs (ExGra-Med, LLaVA-Med) and general-purpose VLMs (Qwen2.5-VL, InternVL2.5), showing that SV-CoT supervision significantly improves interpretability, grounding fidelity, and robustness. Beyond benchmarking, we study its synergy with retrieval-augmented generation, revealing how domain knowledge and visual grounding interact during autoregressive reasoning. Finally, we propose a new mechanism that strengthens the alignment between visual evidence and reasoning, improving both reliability and efficiency. S-Chain establishes a new benchmark for grounded medical reasoning and paves the way toward more trustworthy and explainable medical VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22728
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle S-Chain: Structured Visual Chain-of-Thought For Medicine
Le-Duc, Khai
Nguyen, Duy M. H.
Trinh, Phuong T. H.
Nguyen, Tien-Phat
Diep, Nghiem T.
Ngo, An
Vu, Tung
Vuong, Trinh
Nguyen, Anh-Tien
Nguyen, Mau
Hoang, Van Trung
Nguyen, Khai-Nguyen
Nguyen, Hy
Ngo, Chris
Liu, Anji
Ho, Nhat
Hauschild, Anne-Christin
Nguyen, Khanh Xuan
Nguyen-Tang, Thanh
Xie, Pengtao
Sonntag, Daniel
Zou, James
Niepert, Mathias
Nguyen, Anh Totti
Machine Learning
Computer Vision and Pattern Recognition
Faithful reasoning in medical vision-language models (VLMs) requires not only accurate predictions but also transparent alignment between textual rationales and visual evidence. While Chain-of-Thought (CoT) prompting has shown promise in medical visual question answering (VQA), no large-scale expert-level dataset has captured stepwise reasoning with precise visual grounding. We introduce S-Chain, the first large-scale dataset of 12,000 expert-annotated medical images with bounding boxes and structured visual CoT (SV-CoT), explicitly linking visual regions to reasoning steps. The dataset further supports 16 languages, totaling over 700k VQA pairs for broad multilingual applicability. Using S-Chain, we benchmark state-of-the-art medical VLMs (ExGra-Med, LLaVA-Med) and general-purpose VLMs (Qwen2.5-VL, InternVL2.5), showing that SV-CoT supervision significantly improves interpretability, grounding fidelity, and robustness. Beyond benchmarking, we study its synergy with retrieval-augmented generation, revealing how domain knowledge and visual grounding interact during autoregressive reasoning. Finally, we propose a new mechanism that strengthens the alignment between visual evidence and reasoning, improving both reliability and efficiency. S-Chain establishes a new benchmark for grounded medical reasoning and paves the way toward more trustworthy and explainable medical VLMs.
title S-Chain: Structured Visual Chain-of-Thought For Medicine
topic Machine Learning
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.22728