CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
Fuente:
arXiv
Saved in:
| Main Authors: | Qian, Jiahe, Shen, Yuhao, Chen, Zhangtianyi, Zhou, Juexiao, Wang, Peisong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
by: Kong, Quan, et al.
Published: (2026)
by: Kong, Quan, et al.
Published: (2026)
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
by: Shen, Yuhao, et al.
Published: (2024)
by: Shen, Yuhao, et al.
Published: (2024)
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
by: Chen, Zhangtianyi, et al.
Published: (2026)
by: Chen, Zhangtianyi, et al.
Published: (2026)
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
by: Wu, Jiahe, et al.
Published: (2026)
by: Wu, Jiahe, et al.
Published: (2026)
ForgeryTTT: Zero-Shot Image Manipulation Localization with Test-Time Training
by: Liu, Weihuang, et al.
Published: (2024)
by: Liu, Weihuang, et al.
Published: (2024)
Med-TTT: Vision Test-Time Training model for Medical Image Segmentation
by: Xu, Jiashu
Published: (2024)
by: Xu, Jiashu
Published: (2024)
Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
by: Liu, Fangfu, et al.
Published: (2026)
by: Liu, Fangfu, et al.
Published: (2026)
On the Role of Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
MedCoT: Medical Chain of Thought via Hierarchical Expert
by: Liu, Jiaxiang, et al.
Published: (2024)
by: Liu, Jiaxiang, et al.
Published: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
by: Abdelrahman, Eslam, et al.
Published: (2023)
by: Abdelrahman, Eslam, et al.
Published: (2023)
TTT3R: 3D Reconstruction as Test-Time Training
by: Chen, Xingyu, et al.
Published: (2025)
by: Chen, Xingyu, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
by: Yi, Shixin, et al.
Published: (2025)
by: Yi, Shixin, et al.
Published: (2025)
Trustworthy and Fair SkinGPT-R1 for Democratizing Dermatological Reasoning across Diverse Ethnicities
by: Shen, Yuhao, et al.
Published: (2025)
by: Shen, Yuhao, et al.
Published: (2025)
Measuring Faithful and Plausible Visual Grounding in VQA
by: Reich, Daniel, et al.
Published: (2023)
by: Reich, Daniel, et al.
Published: (2023)
NC-TTT: A Noise Contrastive Approach for Test-Time Training
by: Osowiechi, David, et al.
Published: (2024)
by: Osowiechi, David, et al.
Published: (2024)
ReC-TTT: Contrastive Feature Reconstruction for Test-Time Training
by: Colussi, Marco, et al.
Published: (2024)
by: Colussi, Marco, et al.
Published: (2024)
ClinCoT: Clinical-Aware Visual Chain-of-Thought for Medical Vision Language Models
by: Liu, Xiwei, et al.
Published: (2026)
by: Liu, Xiwei, et al.
Published: (2026)
CoMT: Chain-of-Medical-Thought Reduces Hallucination in Medical Report Generation
by: Jiang, Yue, et al.
Published: (2024)
by: Jiang, Yue, et al.
Published: (2024)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
AU-TTT: Vision Test-Time Training model for Facial Action Unit Detection
by: Xing, Bohao, et al.
Published: (2025)
by: Xing, Bohao, et al.
Published: (2025)
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training
by: Bi, Xiuli, et al.
Published: (2024)
by: Bi, Xiuli, et al.
Published: (2024)
LoRA-TTT: Low-Rank Test-Time Training for Vision-Language Models
by: Kojima, Yuto, et al.
Published: (2025)
by: Kojima, Yuto, et al.
Published: (2025)
Uncovering the Full Potential of Visual Grounding Methods in VQA
by: Reich, Daniel, et al.
Published: (2024)
by: Reich, Daniel, et al.
Published: (2024)
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
by: Chen, Xinyan, et al.
Published: (2025)
by: Chen, Xinyan, et al.
Published: (2025)
DocTTT: Test-Time Training for Handwritten Document Recognition Using Meta-Auxiliary Learning
by: Gu, Wenhao, et al.
Published: (2025)
by: Gu, Wenhao, et al.
Published: (2025)
Decoupled Competitive Framework for Semi-supervised Medical Image Segmentation
by: Chen, Jiahe, et al.
Published: (2025)
by: Chen, Jiahe, et al.
Published: (2025)
PMC-VQA: Visual Instruction Tuning for Medical Visual Question Answering
by: Zhang, Xiaoman, et al.
Published: (2023)
by: Zhang, Xiaoman, et al.
Published: (2023)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
by: Mo, Ye, et al.
Published: (2025)
by: Mo, Ye, et al.
Published: (2025)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Grounded Chain-of-Thought for Multimodal Large Language Models
by: Wu, Qiong, et al.
Published: (2025)
by: Wu, Qiong, et al.
Published: (2025)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
by: Wei, Lai, et al.
Published: (2024)
by: Wei, Lai, et al.
Published: (2024)
TTT-Unet: Enhancing U-Net with Test-Time Training Layers for Biomedical Image Segmentation
by: Zhou, Rong, et al.
Published: (2024)
by: Zhou, Rong, et al.
Published: (2024)
Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
by: Shao, Hao, et al.
Published: (2024)
by: Shao, Hao, et al.
Published: (2024)
Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
by: Etchegaray, Djamahl, et al.
Published: (2025)
by: Etchegaray, Djamahl, et al.
Published: (2025)
Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
by: Safwan, Itbaan, et al.
Published: (2025)
by: Safwan, Itbaan, et al.
Published: (2025)
Pathology-CoT: Learning Visual Chain-of-Thought Agent from Expert Whole Slide Image Diagnosis Behavior
by: Wang, Sheng, et al.
Published: (2025)
by: Wang, Sheng, et al.
Published: (2025)
Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
by: Acuna, David, et al.
Published: (2025)
by: Acuna, David, et al.
Published: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Similar Items
-
Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
by: Shen, Yuhao, et al.
Published: (2025) -
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
by: Kong, Quan, et al.
Published: (2026) -
SkinCaRe: A Multimodal Dermatology Dataset Annotated with Medical Caption and Chain-of-Thought Reasoning
by: Shen, Yuhao, et al.
Published: (2024) -
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
by: Chen, Zhangtianyi, et al.
Published: (2026) -
RGBX-R1: Visual Modality Chain-of-Thought Guided Reinforcement Learning for Multimodal Grounding
by: Wu, Jiahe, et al.
Published: (2026)