BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mullappilly, Sahal Shaji, Kurpath, Mohammed Irfan, Pieri, Sara, Alseiari, Saeed Yahya, Cholakkal, Shanavas, Aldahmani, Khaled, Khan, Fahad, Anwer, Rao, Khan, Salman, Baldwin, Timothy, Cholakkal, Hisham
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915590370754560
author Mullappilly, Sahal Shaji
Kurpath, Mohammed Irfan
Pieri, Sara
Alseiari, Saeed Yahya
Cholakkal, Shanavas
Aldahmani, Khaled
Khan, Fahad
Anwer, Rao
Khan, Salman
Baldwin, Timothy
Cholakkal, Hisham
author_facet Mullappilly, Sahal Shaji
Kurpath, Mohammed Irfan
Pieri, Sara
Alseiari, Saeed Yahya
Cholakkal, Shanavas
Aldahmani, Khaled
Khan, Fahad
Anwer, Rao
Khan, Salman
Baldwin, Timothy
Cholakkal, Hisham
contents We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histology. To train BiMediX2, we curate BiMed-V, an extensive Arabic-English bilingual healthcare dataset consisting of 1.6M samples of diverse medical interactions. This dataset supports a range of medical Large Language Model (LLM) and Large Multimodal Model (LMM) tasks, including multi-turn medical conversations, report generation, and visual question answering (VQA). We also introduce BiMed-MBench, the first Arabic-English medical LMM evaluation benchmark, verified by medical experts. BiMediX2 demonstrates excellent performance across multiple medical LLM and LMM benchmarks, achieving state-of-the-art results compared to other open-sourced models. On BiMed-MBench, BiMediX2 outperforms existing methods by over 9% in English and more than 20% in Arabic evaluations. Additionally, it surpasses GPT-4 by approximately 9% in UPHILL factual accuracy evaluations and excels in various medical VQA, report generation, and report summarization tasks. Our trained models, instruction set, and source code are available at https://github.com/mbzuai-oryx/BiMediX2
format Preprint
id arxiv_https___arxiv_org_abs_2412_07769
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
Mullappilly, Sahal Shaji
Kurpath, Mohammed Irfan
Pieri, Sara
Alseiari, Saeed Yahya
Cholakkal, Shanavas
Aldahmani, Khaled
Khan, Fahad
Anwer, Rao
Khan, Salman
Baldwin, Timothy
Cholakkal, Hisham
Computer Vision and Pattern Recognition
We introduce BiMediX2, a bilingual (Arabic-English) Bio-Medical EXpert Large Multimodal Model that supports text-based and image-based medical interactions. It enables multi-turn conversation in Arabic and English and supports diverse medical imaging modalities, including radiology, CT, and histology. To train BiMediX2, we curate BiMed-V, an extensive Arabic-English bilingual healthcare dataset consisting of 1.6M samples of diverse medical interactions. This dataset supports a range of medical Large Language Model (LLM) and Large Multimodal Model (LMM) tasks, including multi-turn medical conversations, report generation, and visual question answering (VQA). We also introduce BiMed-MBench, the first Arabic-English medical LMM evaluation benchmark, verified by medical experts. BiMediX2 demonstrates excellent performance across multiple medical LLM and LMM benchmarks, achieving state-of-the-art results compared to other open-sourced models. On BiMed-MBench, BiMediX2 outperforms existing methods by over 9% in English and more than 20% in Arabic evaluations. Additionally, it surpasses GPT-4 by approximately 9% in UPHILL factual accuracy evaluations and excels in various medical VQA, report generation, and report summarization tasks. Our trained models, instruction set, and source code are available at https://github.com/mbzuai-oryx/BiMediX2
title BiMediX2: Bio-Medical EXpert LMM for Diverse Medical Modalities
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.07769