MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Lijian, Sun, Hao, Ni, Ziyu, Li, Hongsheng, Zhang, Shaoting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A foundation model for generalizable disease diagnosis in chest X-ray images
von: Xu, Lijian, et al.
Veröffentlicht: (2024)
von: Xu, Lijian, et al.
Veröffentlicht: (2024)
Learning A Multi-Task Transformer Via Unified And Customized Instruction Tuning For Chest Radiograph Interpretation
von: Xu, Lijian, et al.
Veröffentlicht: (2023)
von: Xu, Lijian, et al.
Veröffentlicht: (2023)
Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023)
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024)
DeViDe: Faceted medical knowledge for improved medical vision-language pre-training
von: Luo, Haozhe, et al.
Veröffentlicht: (2024)
von: Luo, Haozhe, et al.
Veröffentlicht: (2024)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
von: Wu, Yiqi, et al.
Veröffentlicht: (2024)
von: Wu, Yiqi, et al.
Veröffentlicht: (2024)
MedGEN-Bench: Contextually entangled benchmark for open-ended multimodal medical generation
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
Bridging visual saliency and large language models for explainable deep learning in medical imaging
von: Nguezet, Paul Valery, et al.
Veröffentlicht: (2026)
von: Nguezet, Paul Valery, et al.
Veröffentlicht: (2026)
A multimodal vision foundation model for generalizable knee pathology
von: Yu, Kang, et al.
Veröffentlicht: (2026)
von: Yu, Kang, et al.
Veröffentlicht: (2026)
A generalizable foundation model for intraoperative understanding across surgical procedures
von: Park, Kanggil, et al.
Veröffentlicht: (2026)
von: Park, Kanggil, et al.
Veröffentlicht: (2026)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
ViSTa Dataset: Do vision-language models understand sequential tasks?
von: Wybitul, Evžen, et al.
Veröffentlicht: (2024)
von: Wybitul, Evžen, et al.
Veröffentlicht: (2024)
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
von: Xing, Yang, et al.
Veröffentlicht: (2026)
von: Xing, Yang, et al.
Veröffentlicht: (2026)
A generalizable large-scale foundation model for musculoskeletal radiographs
von: Kim, Shinn, et al.
Veröffentlicht: (2026)
von: Kim, Shinn, et al.
Veröffentlicht: (2026)
Visual concept ranking uncovers medical shortcuts used by large multimodal models
von: Janizek, Joseph D., et al.
Veröffentlicht: (2026)
von: Janizek, Joseph D., et al.
Veröffentlicht: (2026)
UCell: rethinking generalizability and scaling of bio-medical vision models
von: Kuang, Nicholas, et al.
Veröffentlicht: (2026)
von: Kuang, Nicholas, et al.
Veröffentlicht: (2026)
On the robustness of multimodal language model towards distractions
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
A multi-modal vision-language model for generalizable annotation-free pathology localization
von: Yang, Hao, et al.
Veröffentlicht: (2024)
von: Yang, Hao, et al.
Veröffentlicht: (2024)
Do large language vision models understand 3D shapes?
von: Eppel, Sagi
Veröffentlicht: (2024)
von: Eppel, Sagi
Veröffentlicht: (2024)
SalsaAgent: A multimodal embodied language model for interactive dance generation
von: Yazdian, Payam Jome, et al.
Veröffentlicht: (2026)
von: Yazdian, Payam Jome, et al.
Veröffentlicht: (2026)
Evaluating point-light biological motion in multimodal large language models
von: Kadambi, Akila, et al.
Veröffentlicht: (2025)
von: Kadambi, Akila, et al.
Veröffentlicht: (2025)
In-context learning enables multimodal large language models to classify cancer pathology images
von: Ferber, Dyke, et al.
Veröffentlicht: (2024)
von: Ferber, Dyke, et al.
Veröffentlicht: (2024)
A generalizable 3D framework and model for self-supervised learning in medical imaging
von: Xu, Tony, et al.
Veröffentlicht: (2025)
von: Xu, Tony, et al.
Veröffentlicht: (2025)
LLaVAction: evaluating and training multi-modal large language models for action understanding
von: Qi, Haozhe, et al.
Veröffentlicht: (2025)
von: Qi, Haozhe, et al.
Veröffentlicht: (2025)
Scaling medical imaging report generation with multimodal reinforcement learning
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
von: Liu, Qianchu, et al.
Veröffentlicht: (2026)
A benchmark multimodal oro-dental dataset for large vision-language models
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
von: Pang, Yik Lung, et al.
Veröffentlicht: (2026)
von: Pang, Yik Lung, et al.
Veröffentlicht: (2026)
Explaining latent representations of generative models with large multimodal models
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
von: Zhu, Mengdan, et al.
Veröffentlicht: (2024)
MedLSAM: Localize and Segment Anything Model for 3D CT Images
von: Lei, Wenhui, et al.
Veröffentlicht: (2023)
von: Lei, Wenhui, et al.
Veröffentlicht: (2023)
Explainable artificial intelligence (XAI): from inherent explainability to large language models
von: Mumuni, Fuseini, et al.
Veröffentlicht: (2025)
von: Mumuni, Fuseini, et al.
Veröffentlicht: (2025)
Mobile U-ViT: Revisiting large kernel and U-shaped ViT for efficient medical image segmentation
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
von: Tang, Fenghe, et al.
Veröffentlicht: (2025)
ProtoS-ViT: Visual foundation models for sparse self-explainable classifications
von: Turbé, Hugues, et al.
Veröffentlicht: (2024)
von: Turbé, Hugues, et al.
Veröffentlicht: (2024)
Comprehensive language-image pre-training for 3D medical image understanding
von: Wald, Tassilo, et al.
Veröffentlicht: (2025)
von: Wald, Tassilo, et al.
Veröffentlicht: (2025)
MedDiff-FM: A Diffusion-based Foundation Model for Versatile Medical Image Applications
von: Yu, Yongrui, et al.
Veröffentlicht: (2024)
von: Yu, Yongrui, et al.
Veröffentlicht: (2024)
MAIRA-1: A specialised large multimodal model for radiology report generation
von: Hyland, Stephanie L., et al.
Veröffentlicht: (2023)
von: Hyland, Stephanie L., et al.
Veröffentlicht: (2023)
MedIAnomaly: A comparative study of anomaly detection in medical images
von: Cai, Yu, et al.
Veröffentlicht: (2024)
von: Cai, Yu, et al.
Veröffentlicht: (2024)
MedCAL-Bench: A Comprehensive Benchmark on Cold-Start Active Learning with Foundation Models for Medical Image Analysis
von: Zhu, Ning, et al.
Veröffentlicht: (2025)
von: Zhu, Ning, et al.
Veröffentlicht: (2025)
Elucidating the design space of language models for image generation
von: Liu, Xuantong, et al.
Veröffentlicht: (2024)
von: Liu, Xuantong, et al.
Veröffentlicht: (2024)
TinyViM: Frequency Decoupling for Tiny Hybrid Vision Mamba
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
von: Ma, Xiaowen, et al.
Veröffentlicht: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
A foundation model for generalizable disease diagnosis in chest X-ray images
von: Xu, Lijian, et al.
Veröffentlicht: (2024) -
Learning A Multi-Task Transformer Via Unified And Customized Instruction Tuning For Chest Radiograph Interpretation
von: Xu, Lijian, et al.
Veröffentlicht: (2023) -
Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2023) -
One Leaf Reveals the Season: Occlusion-Based Contrastive Learning with Semantic-Aware Views for Efficient Visual Representation
von: Yang, Xiaoyu, et al.
Veröffentlicht: (2024) -
DeViDe: Faceted medical knowledge for improved medical vision-language pre-training
von: Luo, Haozhe, et al.
Veröffentlicht: (2024)