Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Kaitao, Rui, Shaohao, Jiang, Yankai, Wu, Jiamin, Zheng, Qihao, Song, Chunfeng, Wang, Xiaosong, Zhou, Mu, Liu, Mianxin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
di: Chen, Kaitao, et al.
Pubblicazione: (2025)
di: Chen, Kaitao, et al.
Pubblicazione: (2025)
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
Multi-modal Vision Pre-training for Medical Image Analysis
di: Rui, Shaohao, et al.
Pubblicazione: (2024)
di: Rui, Shaohao, et al.
Pubblicazione: (2024)
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
di: Li, Yiwei, et al.
Pubblicazione: (2026)
di: Li, Yiwei, et al.
Pubblicazione: (2026)
CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
di: Rui, Shaohao, et al.
Pubblicazione: (2025)
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
di: Guo, Zhanqiang, et al.
Pubblicazione: (2024)
di: Guo, Zhanqiang, et al.
Pubblicazione: (2024)
InvCoSS: Inversion-driven Continual Self-supervised Learning in Medical Multi-modal Image Pre-training
di: Luo, Zihao, et al.
Pubblicazione: (2025)
di: Luo, Zihao, et al.
Pubblicazione: (2025)
Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
di: Sun, Haoran, et al.
Pubblicazione: (2025)
di: Sun, Haoran, et al.
Pubblicazione: (2025)
CTSL: Codebook-based Temporal-Spatial Learning for Accurate Non-Contrast Cardiac Risk Prediction Using Cine MRIs
di: Su, Haoyang, et al.
Pubblicazione: (2025)
di: Su, Haoyang, et al.
Pubblicazione: (2025)
Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
di: Qin, Yiming, et al.
Pubblicazione: (2025)
di: Qin, Yiming, et al.
Pubblicazione: (2025)
To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs
di: Hong, Rui, et al.
Pubblicazione: (2026)
di: Hong, Rui, et al.
Pubblicazione: (2026)
SynBrain: Enhancing Visual-to-fMRI Synthesis via Probabilistic Representation Learning
di: Mai, Weijian, et al.
Pubblicazione: (2025)
di: Mai, Weijian, et al.
Pubblicazione: (2025)
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
di: Liu, Chengzhi, et al.
Pubblicazione: (2025)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
di: Tang, Fei, et al.
Pubblicazione: (2025)
di: Tang, Fei, et al.
Pubblicazione: (2025)
MindAligner: Explicit Brain Functional Alignment for Cross-Subject Visual Decoding from Limited fMRI Data
di: Dai, Yuqin, et al.
Pubblicazione: (2025)
di: Dai, Yuqin, et al.
Pubblicazione: (2025)
Cost-effective Instruction Learning for Pathology Vision and Language Analysis
di: Chen, Kaitao, et al.
Pubblicazione: (2024)
di: Chen, Kaitao, et al.
Pubblicazione: (2024)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
di: Liu, Shuming, et al.
Pubblicazione: (2026)
di: Liu, Shuming, et al.
Pubblicazione: (2026)
Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examination
di: Shi, Chufan, et al.
Pubblicazione: (2026)
di: Shi, Chufan, et al.
Pubblicazione: (2026)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
di: Zhang, Yongchang, et al.
Pubblicazione: (2026)
di: Zhang, Yongchang, et al.
Pubblicazione: (2026)
BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
di: Ye, Junyan, et al.
Pubblicazione: (2025)
di: Ye, Junyan, et al.
Pubblicazione: (2025)
Language-Enhanced Generative Modeling for Amyloid PET Synthesis from MRI and Blood Biomarkers
di: Zhang, Zhengjie, et al.
Pubblicazione: (2025)
di: Zhang, Zhengjie, et al.
Pubblicazione: (2025)
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding
di: Yang, Jianjiang, et al.
Pubblicazione: (2025)
di: Yang, Jianjiang, et al.
Pubblicazione: (2025)
VisualActBench: Can VLMs See and Act like a Human?
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
di: Zhang, Daoan, et al.
Pubblicazione: (2025)
Right this way: Can VLMs Guide Us to See More to Answer Questions?
di: Liu, Li, et al.
Pubblicazione: (2024)
di: Liu, Li, et al.
Pubblicazione: (2024)
AdaBrain-Bench: Benchmarking Brain Foundation Models for Brain-Computer Interface Applications
di: Wu, Jiamin, et al.
Pubblicazione: (2025)
di: Wu, Jiamin, et al.
Pubblicazione: (2025)
Towards All-in-One Medical Image Re-Identification
di: Tian, Yuan, et al.
Pubblicazione: (2025)
di: Tian, Yuan, et al.
Pubblicazione: (2025)
PRISM: A Framework Harnessing Unsupervised Visual Representations and Textual Prompts for Explainable MACE Survival Prediction from Cardiac Cine MRI
di: Su, Haoyang, et al.
Pubblicazione: (2025)
di: Su, Haoyang, et al.
Pubblicazione: (2025)
Think Twice Before Selection: Federated Evidential Active Learning for Medical Image Analysis with Domain Shifts
di: Chen, Jiayi, et al.
Pubblicazione: (2023)
di: Chen, Jiayi, et al.
Pubblicazione: (2023)
See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
di: Zhang, Shuoshuo, et al.
Pubblicazione: (2025)
di: Zhang, Shuoshuo, et al.
Pubblicazione: (2025)
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning
di: Zeng, Fanhu, et al.
Pubblicazione: (2026)
di: Zeng, Fanhu, et al.
Pubblicazione: (2026)
VITAL: Visual-Semantic Dual Supervision for Enhanced and Interpretable Latent Reasoning in Medical MLLMs
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
di: Li, Qiaoru, et al.
Pubblicazione: (2026)
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection
di: Wu, Zhiheng, et al.
Pubblicazione: (2026)
di: Wu, Zhiheng, et al.
Pubblicazione: (2026)
Incentivizing Tool-augmented Thinking with Images for Medical Image Analysis
di: Jiang, Yankai, et al.
Pubblicazione: (2025)
di: Jiang, Yankai, et al.
Pubblicazione: (2025)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
di: Dai, Ziyun, et al.
Pubblicazione: (2025)
di: Dai, Ziyun, et al.
Pubblicazione: (2025)
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
di: Liu, Zhining, et al.
Pubblicazione: (2025)
di: Liu, Zhining, et al.
Pubblicazione: (2025)
VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
di: Berman, Shmuel, et al.
Pubblicazione: (2025)
TK-Mamba: Marrying KAN With Mamba for Text-Driven 3D Medical Image Segmentation
di: Yang, Haoyu, et al.
Pubblicazione: (2025)
di: Yang, Haoyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
di: Chen, Kaitao, et al.
Pubblicazione: (2025) -
AdaThink-Med: Medical Adaptive Thinking with Uncertainty-Guided Length Calibration
di: Rui, Shaohao, et al.
Pubblicazione: (2025) -
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
di: Rui, Shaohao, et al.
Pubblicazione: (2025) -
Multi-modal Vision Pre-training for Medical Image Analysis
di: Rui, Shaohao, et al.
Pubblicazione: (2024) -
Thinking with Gaze: Sequential Eye-Tracking as Visual Reasoning Supervision for Medical VLMs
di: Li, Yiwei, et al.
Pubblicazione: (2026)