ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Nguyen, Duy M. H., Diep, Nghiem T., Nguyen, Trung Q., Le, Hoang-Bao, Nguyen, Tai, Nguyen, Tien, Nguyen, TrungTin, Ho, Nhat, Xie, Pengtao, Wattenhofer, Roger, Zou, James, Sonntag, Daniel, Niepert, Mathias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918189407928320
author Nguyen, Duy M. H.
Diep, Nghiem T.
Nguyen, Trung Q.
Le, Hoang-Bao
Nguyen, Tai
Nguyen, Tien
Nguyen, TrungTin
Ho, Nhat
Xie, Pengtao
Wattenhofer, Roger
Zou, James
Sonntag, Daniel
Niepert, Mathias
author_facet Nguyen, Duy M. H.
Diep, Nghiem T.
Nguyen, Trung Q.
Le, Hoang-Bao
Nguyen, Tai
Nguyen, Tien
Nguyen, TrungTin
Ho, Nhat
Xie, Pengtao
Wattenhofer, Roger
Zou, James
Sonntag, Daniel
Niepert, Mathias
contents State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMA-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med's performance using just 10% of the pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI.
format Preprint
id arxiv_https___arxiv_org_abs_2410_02615
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
Nguyen, Duy M. H.
Diep, Nghiem T.
Nguyen, Trung Q.
Le, Hoang-Bao
Nguyen, Tai
Nguyen, Tien
Nguyen, TrungTin
Ho, Nhat
Xie, Pengtao
Wattenhofer, Roger
Zou, James
Sonntag, Daniel
Niepert, Mathias
Machine Learning
State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMA-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med's performance using just 10% of the pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI.
title ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
topic Machine Learning
url https://arxiv.org/abs/2410.02615