ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918189407928320 |
|---|---|
| author | Nguyen, Duy M. H. Diep, Nghiem T. Nguyen, Trung Q. Le, Hoang-Bao Nguyen, Tai Nguyen, Tien Nguyen, TrungTin Ho, Nhat Xie, Pengtao Wattenhofer, Roger Zou, James Sonntag, Daniel Niepert, Mathias |
| author_facet | Nguyen, Duy M. H. Diep, Nghiem T. Nguyen, Trung Q. Le, Hoang-Bao Nguyen, Tai Nguyen, Tien Nguyen, TrungTin Ho, Nhat Xie, Pengtao Wattenhofer, Roger Zou, James Sonntag, Daniel Niepert, Mathias |
| contents | State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMA-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med's performance using just 10% of the pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2410_02615 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models Nguyen, Duy M. H. Diep, Nghiem T. Nguyen, Trung Q. Le, Hoang-Bao Nguyen, Tai Nguyen, Tien Nguyen, TrungTin Ho, Nhat Xie, Pengtao Wattenhofer, Roger Zou, James Sonntag, Daniel Niepert, Mathias Machine Learning State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMA-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med's performance using just 10% of the pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI. |
| title | ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models |
| topic | Machine Learning |
| url | https://arxiv.org/abs/2410.02615 |