Making medical vision-language models think causally across modalities with retrieval-augmented cross-modal reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Weiqin, Xue, Haowen, Peng, Qingyi, Hu, Hexuan, Huang, Qian, Zhang, Tingbo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Training-free retrieval-augmented generation with reinforced reasoning for flood damage nowcasting
by: Huang, Lipai, et al.
Published: (2026)
by: Huang, Lipai, et al.
Published: (2026)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025)
by: Rainone, Corrado, et al.
Published: (2025)
CaseGPT: a case reasoning framework based on language models and retrieval-augmented generation
by: Yang, Rui
Published: (2024)
by: Yang, Rui
Published: (2024)
Agentic retrieval-augmented reasoning reshapes collective reliability under model variability in radiology question answering
by: Farajiamiri, Mina, et al.
Published: (2026)
by: Farajiamiri, Mina, et al.
Published: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
by: Liu, Shuliang, et al.
Published: (2025)
by: Liu, Shuliang, et al.
Published: (2025)
Multi-step retrieval and reasoning improves radiology question answering with large language models
by: Wind, Sebastian, et al.
Published: (2025)
by: Wind, Sebastian, et al.
Published: (2025)
Multi-modal clothing recommendation model based on large model and VAE enhancement
by: Huang, Bingjie, et al.
Published: (2024)
by: Huang, Bingjie, et al.
Published: (2024)
Contrastive-learning of language embedding and biological features for cross modality encoding and effector prediction.
by: Peng, Yue, et al.
Published: (2025)
by: Peng, Yue, et al.
Published: (2025)
AGA: An adaptive group alignment framework for structured medical cross-modal representation learning
by: Li, Wei, et al.
Published: (2025)
by: Li, Wei, et al.
Published: (2025)
Robust image classification with multi-modal large language models
by: Villani, Francesco, et al.
Published: (2024)
by: Villani, Francesco, et al.
Published: (2024)
Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
by: Li, Haobin, et al.
Published: (2025)
by: Li, Haobin, et al.
Published: (2025)
Test-time Adaptation for Cross-modal Retrieval with Query Shift
by: Li, Haobin, et al.
Published: (2024)
by: Li, Haobin, et al.
Published: (2024)
Promoting cross-modal representations to improve multimodal foundation models for physiological signals
by: Fang, Ching, et al.
Published: (2024)
by: Fang, Ching, et al.
Published: (2024)
Autosen: improving automatic wifi human sensing through cross-modal autoencoder
by: Gao, Qian, et al.
Published: (2024)
by: Gao, Qian, et al.
Published: (2024)
Retrieval-augmented reasoning with lean language models
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
by: Chan, Ryan Sze-Yin, et al.
Published: (2025)
On the Comparison between Multi-modal and Single-modal Contrastive Learning
by: Huang, Wei, et al.
Published: (2024)
by: Huang, Wei, et al.
Published: (2024)
Cross-modal Active Complementary Learning with Self-refining Correspondence
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
Toward Robust and Harmonious Adaptation for Cross-modal Retrieval
by: Li, Haobin, et al.
Published: (2025)
by: Li, Haobin, et al.
Published: (2025)
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
Multi-modal Co-learning for Earth Observation: Enhancing single-modality models via modality collaboration
by: Mena, Francisco, et al.
Published: (2025)
by: Mena, Francisco, et al.
Published: (2025)
A multi-modal vision-language model for generalizable annotation-free pathology localization
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Understanding protein function with a multimodal retrieval-augmented foundation model
by: Truong Jr, Timothy Fei, et al.
Published: (2025)
by: Truong Jr, Timothy Fei, et al.
Published: (2025)
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay
by: Oota, Subba Reddy, et al.
Published: (2026)
by: Oota, Subba Reddy, et al.
Published: (2026)
Multi-modal brain encoding models for multi-modal stimuli
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
CAREL: Instruction-guided reinforcement learning with cross-modal auxiliary objectives
by: Saghafian, Armin, et al.
Published: (2024)
by: Saghafian, Armin, et al.
Published: (2024)
Learnable Cross-modal Knowledge Distillation for Multi-modal Learning with Missing Modality
by: Wang, Hu, et al.
Published: (2023)
by: Wang, Hu, et al.
Published: (2023)
FiMMIA: scaling semantic perturbation-based membership inference across modalities
by: Emelyanov, Anton, et al.
Published: (2025)
by: Emelyanov, Anton, et al.
Published: (2025)
Balance-aware Sequence Sampling Makes Multi-modal Learning Better
by: Guan, Zhi-Hao
Published: (2025)
by: Guan, Zhi-Hao
Published: (2025)
bi-modal textual prompt learning for vision-language models in remote sensing
by: Kashyap, Pankhi, et al.
Published: (2026)
by: Kashyap, Pankhi, et al.
Published: (2026)
Relational reasoning and inductive bias in transformers and large language models
by: Geerts, Jesse, et al.
Published: (2025)
by: Geerts, Jesse, et al.
Published: (2025)
What explains the success of cross-modal fine-tuning with ORCA?
by: García-de-Herreros, Paloma, et al.
Published: (2024)
by: García-de-Herreros, Paloma, et al.
Published: (2024)
Understanding the Robustness of Multi-modal Contrastive Learning to Distribution Shift
by: Xue, Yihao, et al.
Published: (2023)
by: Xue, Yihao, et al.
Published: (2023)
LLM-Cave: A benchmark and light environment for large language models reasoning and decision-making system
by: Li, Huanyu, et al.
Published: (2025)
by: Li, Huanyu, et al.
Published: (2025)
Contextualization of ASR with LLM using phonetic retrieval-based augmentation
by: Lei, Zhihong, et al.
Published: (2024)
by: Lei, Zhihong, et al.
Published: (2024)
An analysis of vision-language models for fabric retrieval
by: Giuliari, Francesco, et al.
Published: (2025)
by: Giuliari, Francesco, et al.
Published: (2025)
Look, Remember and Reason: Grounded reasoning in videos with language models
by: Bhattacharyya, Apratim, et al.
Published: (2023)
by: Bhattacharyya, Apratim, et al.
Published: (2023)
Learning to reason about rare diseases through retrieval-augmented agents
by: Kim, Ha Young, et al.
Published: (2025)
by: Kim, Ha Young, et al.
Published: (2025)
A foundation model enpowered by a multi-modal prompt engine for universal seismic geobody interpretation across surveys
by: Gao, Hang, et al.
Published: (2024)
by: Gao, Hang, et al.
Published: (2024)
Similar Items
-
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025) -
Training-free retrieval-augmented generation with reinforced reasoning for flood damage nowcasting
by: Huang, Lipai, et al.
Published: (2026) -
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026) -
Replacing thinking with tool usage enables reasoning in small language models
by: Rainone, Corrado, et al.
Published: (2025) -
CaseGPT: a case reasoning framework based on language models and retrieval-augmented generation
by: Yang, Rui
Published: (2024)