MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Yanyuan, Xu, Dexuan, Huang, Yu, Zhan, Songkun, Wang, Hanpin, Chen, Dongxue, Wang, Xueping, Qiu, Meikang, Li, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024)
by: Matos, João, et al.
Published: (2024)
Protecting multimodal large language models against misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025)
by: Zhang, Ruixuan, et al.
Published: (2025)
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
by: Oota, Subba Reddy, et al.
Published: (2025)
by: Oota, Subba Reddy, et al.
Published: (2025)
MARCUS: An agentic, multimodal vision-language model for cardiac diagnosis and management
by: O'Sullivan, Jack W, et al.
Published: (2026)
by: O'Sullivan, Jack W, et al.
Published: (2026)
Towards deployment-centric multimodal AI beyond vision and language
by: Liu, Xianyuan, et al.
Published: (2025)
by: Liu, Xianyuan, et al.
Published: (2025)
Embodiment in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
by: Lv, Haoxin, et al.
Published: (2025)
by: Lv, Haoxin, et al.
Published: (2025)
Revisiting Label Inference Attacks in Vertical Federated Learning: Why They Are Vulnerable and How to Defend
by: Liu, Yige, et al.
Published: (2026)
by: Liu, Yige, et al.
Published: (2026)
Are vision language models robust to uncertain inputs?
by: Wang, Xi, et al.
Published: (2025)
by: Wang, Xi, et al.
Published: (2025)
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
by: Pang, Yik Lung, et al.
Published: (2026)
by: Pang, Yik Lung, et al.
Published: (2026)
Retrieval-augmented in-context learning for multimodal large language models in disease classification
by: Zhan, Zaifu, et al.
Published: (2025)
by: Zhan, Zaifu, et al.
Published: (2025)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
What do vision-language models see in the context? Investigating multimodal in-context learning
by: Santos, Gabriel O. dos, et al.
Published: (2025)
by: Santos, Gabriel O. dos, et al.
Published: (2025)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Visual cognition in multimodal large language models
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
by: Buschoff, Luca M. Schulze, et al.
Published: (2023)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
by: Xing, Yang, et al.
Published: (2026)
by: Xing, Yang, et al.
Published: (2026)
Is your multimodal large language model a good science tutor?
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
Optimal input excitations for suppressing nonlinear instabilities in multimode fibers
by: Wisal, Kabish, et al.
Published: (2024)
by: Wisal, Kabish, et al.
Published: (2024)
Generative vector search to improve pathology foundation models across multimodal vision-language tasks
by: Ekvall, Markus, et al.
Published: (2025)
by: Ekvall, Markus, et al.
Published: (2025)
Combining inherent knowledge of vision-language models with unsupervised domain adaptation through strong-weak guidance
by: Westfechtel, Thomas, et al.
Published: (2023)
by: Westfechtel, Thomas, et al.
Published: (2023)
MedMKEB: A Comprehensive Knowledge Editing Benchmark for Medical Multimodal Large Language Models
by: Xu, Dexuan, et al.
Published: (2025)
by: Xu, Dexuan, et al.
Published: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Human-like object concept representations emerge naturally in multimodal large language models
by: Du, Changde, et al.
Published: (2024)
by: Du, Changde, et al.
Published: (2024)
Exploiting spacetime symmetry in dissipative nonlinear multimode amplifiers for output control
by: Chen, Chun-Wei, et al.
Published: (2024)
by: Chen, Chun-Wei, et al.
Published: (2024)
Probing the limitations of multimodal language models for chemistry and materials research
by: Alampara, Nawaf, et al.
Published: (2024)
by: Alampara, Nawaf, et al.
Published: (2024)
Closing the gap in multimodal medical representation alignment
by: Grassucci, Eleonora, et al.
Published: (2026)
by: Grassucci, Eleonora, et al.
Published: (2026)
Animalbooth: multimodal feature enhancement for animal subject personalization
by: Liu, Chen, et al.
Published: (2025)
by: Liu, Chen, et al.
Published: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
New frontiers in artificial intelligence for biodiversity research and conservation with multimodal language models
by: Zhongqi Miao, et al.
Published: (2025)
by: Zhongqi Miao, et al.
Published: (2025)
Vibe-Eval: A hard evaluation suite for measuring progress of multimodal language models
by: Padlewski, Piotr, et al.
Published: (2024)
by: Padlewski, Piotr, et al.
Published: (2024)
Evaluating point-light biological motion in multimodal large language models
by: Kadambi, Akila, et al.
Published: (2025)
by: Kadambi, Akila, et al.
Published: (2025)
DocLLM: A layout-aware generative language model for multimodal document understanding
by: Wang, Dongsheng, et al.
Published: (2023)
by: Wang, Dongsheng, et al.
Published: (2023)
Attacks on multimodal models
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
by: Xu, Lijian, et al.
Published: (2024)
by: Xu, Lijian, et al.
Published: (2024)
YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception
by: Impraimakis, Marios, et al.
Published: (2026)
by: Impraimakis, Marios, et al.
Published: (2026)
MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
by: Janizek, Joseph D., et al.
Published: (2026)
by: Janizek, Joseph D., et al.
Published: (2026)
OmDet: Large-scale vision-language multi-dataset pre-training with multimodal detection network
by: Zhao, Tiancheng, et al.
Published: (2022)
by: Zhao, Tiancheng, et al.
Published: (2022)
Similar Items
-
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025) -
WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation
by: Matos, João, et al.
Published: (2024) -
Protecting multimodal large language models against misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025) -
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
by: Zhang, Ruixuan, et al.
Published: (2025) -
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
by: Oota, Subba Reddy, et al.
Published: (2025)