CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qian, Jiahe, Shen, Yuhao, Chen, Zhangtianyi, Zhou, Juexiao, Wang, Peisong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912712440676352
author Qian, Jiahe
Shen, Yuhao
Chen, Zhangtianyi
Zhou, Juexiao
Wang, Peisong
author_facet Qian, Jiahe
Shen, Yuhao
Chen, Zhangtianyi
Zhou, Juexiao
Wang, Peisong
contents Medical visual question answering could support clinical decision making, yet current systems often fail under domain shift and produce answers that are weakly grounded in image evidence. This reliability gap arises when models attend to spurious regions and when retraining or additional labels are impractical at deployment time. We address this setting with CoTBox-TTT, an evidence-first test-time training approach that adapts a vision-language model at inference while keeping all backbones frozen. The method updates only a small set of continuous soft prompts. It identifies question-relevant regions through a visual chain-of-thought signal and encourages answer consistency across the original image and a localized crop. The procedure is label free, and plug and play with diverse backbones. Experiments on medical VQA show that the approach is practical for real deployments. For instance, adding CoTBox-TTT to LLaVA increases closed-ended accuracy by 12.3% on pathVQA.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12446
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
Qian, Jiahe
Shen, Yuhao
Chen, Zhangtianyi
Zhou, Juexiao
Wang, Peisong
Computer Vision and Pattern Recognition
Medical visual question answering could support clinical decision making, yet current systems often fail under domain shift and produce answers that are weakly grounded in image evidence. This reliability gap arises when models attend to spurious regions and when retraining or additional labels are impractical at deployment time. We address this setting with CoTBox-TTT, an evidence-first test-time training approach that adapts a vision-language model at inference while keeping all backbones frozen. The method updates only a small set of continuous soft prompts. It identifies question-relevant regions through a visual chain-of-thought signal and encourages answer consistency across the original image and a localized crop. The procedure is label free, and plug and play with diverse backbones. Experiments on medical VQA show that the approach is practical for real deployments. For instance, adding CoTBox-TTT to LLaVA increases closed-ended accuracy by 12.3% on pathVQA.
title CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.12446