Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cai, Zhenyang, Chen, Junying, Wang, Rongsheng, Wang, Weihong, Deng, Yonglin, Song, Dingjie, Chen, Yize, Zhang, Zixu, Wang, Benyou
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909630059249664
author Cai, Zhenyang
Chen, Junying
Wang, Rongsheng
Wang, Weihong
Deng, Yonglin
Song, Dingjie
Chen, Yize
Zhang, Zixu
Wang, Benyou
author_facet Cai, Zhenyang
Chen, Junying
Wang, Rongsheng
Wang, Weihong
Deng, Yonglin
Song, Dingjie
Chen, Yize
Zhang, Zixu
Wang, Benyou
contents Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors driving this generalization remain unclear. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG), which refers to the models' ability to understand novel combinations by recombining learned elements, as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and confirmed that MLLMs can achieve CG across classification and detection tasks, underscoring its broader generalization potential. Med-MAT is available at https://github.com/FreedomIntelligence/Med-MAT.
format Preprint
id arxiv_https___arxiv_org_abs_2412_20070
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
Cai, Zhenyang
Chen, Junying
Wang, Rongsheng
Wang, Weihong
Deng, Yonglin
Song, Dingjie
Chen, Yize
Zhang, Zixu
Wang, Benyou
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
Medical imaging provides essential visual insights for diagnosis, and multimodal large language models (MLLMs) are increasingly utilized for its analysis due to their strong generalization capabilities; however, the underlying factors driving this generalization remain unclear. Current research suggests that multi-task training outperforms single-task as different tasks can benefit each other, but they often overlook the internal relationships within these tasks. To analyze this phenomenon, we attempted to employ compositional generalization (CG), which refers to the models' ability to understand novel combinations by recombining learned elements, as a guiding framework. Since medical images can be precisely defined by Modality, Anatomical area, and Task, naturally providing an environment for exploring CG, we assembled 106 medical datasets to create Med-MAT for comprehensive experiments. The experiments confirmed that MLLMs can use CG to understand unseen medical images and identified CG as one of the main drivers of the generalization observed in multi-task training. Additionally, further studies demonstrated that CG effectively supports datasets with limited data and confirmed that MLLMs can achieve CG across classification and detection tasks, underscoring its broader generalization potential. Med-MAT is available at https://github.com/FreedomIntelligence/Med-MAT.
title Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Machine Learning
url https://arxiv.org/abs/2412.20070