LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qin, Zhenyue, Liu, Yang, Yin, Yu, Ding, Jinyu, Zhang, Haoran, Li, Anran, Campbell, Dylan, Wu, Xuansheng, Zou, Ke, Keenan, Tiarnan D. L., Chew, Emily Y., Lu, Zhiyong, Tham, Yih Chung, Liu, Ninghao, Zhang, Xiuzhen, Chen, Qingyu
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915870396121088
author Qin, Zhenyue
Liu, Yang
Yin, Yu
Ding, Jinyu
Zhang, Haoran
Li, Anran
Campbell, Dylan
Wu, Xuansheng
Zou, Ke
Keenan, Tiarnan D. L.
Chew, Emily Y.
Lu, Zhiyong
Tham, Yih Chung
Liu, Ninghao
Zhang, Xiuzhen
Chen, Qingyu
author_facet Qin, Zhenyue
Liu, Yang
Yin, Yu
Ding, Jinyu
Zhang, Haoran
Li, Anran
Campbell, Dylan
Wu, Xuansheng
Zou, Ke
Keenan, Tiarnan D. L.
Chew, Emily Y.
Lu, Zhiyong
Tham, Yih Chung
Liu, Ninghao
Zhang, Xiuzhen
Chen, Qingyu
contents Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image interpretation, advancing MLLMs for ophthalmology is hindered by the lack of comprehensive benchmark datasets suitable for evaluating generative models. We present a large-scale multimodal ophthalmology benchmark comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations, supporting anatomical structure recognition, disease screening, disease staging, and demographic prediction for bias evaluation. This work extends our preliminary LMOD benchmark with three major enhancements: (1) nearly 50% dataset expansion with substantial enlargement of color fundus photography; (2) broadened task coverage including binary disease diagnosis, multi-class diagnosis, severity classification with international grading standards, and demographic prediction; and (3) systematic evaluation of 24 state-of-the-art MLLMs. Our evaluations reveal both promise and limitations. Top-performing models achieved ~58% accuracy in disease screening under zero-shot settings, and performance remained suboptimal for challenging tasks like disease staging. We will publicly release the dataset, curation pipeline, and leaderboard to potentially advance ophthalmic AI applications and reduce the global burden of vision-threatening diseases.
format Preprint
id arxiv_https___arxiv_org_abs_2509_25620
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
Qin, Zhenyue
Liu, Yang
Yin, Yu
Ding, Jinyu
Zhang, Haoran
Li, Anran
Campbell, Dylan
Wu, Xuansheng
Zou, Ke
Keenan, Tiarnan D. L.
Chew, Emily Y.
Lu, Zhiyong
Tham, Yih Chung
Liu, Ninghao
Zhang, Xiuzhen
Chen, Qingyu
Computer Vision and Pattern Recognition
Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image interpretation, advancing MLLMs for ophthalmology is hindered by the lack of comprehensive benchmark datasets suitable for evaluating generative models. We present a large-scale multimodal ophthalmology benchmark comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations, supporting anatomical structure recognition, disease screening, disease staging, and demographic prediction for bias evaluation. This work extends our preliminary LMOD benchmark with three major enhancements: (1) nearly 50% dataset expansion with substantial enlargement of color fundus photography; (2) broadened task coverage including binary disease diagnosis, multi-class diagnosis, severity classification with international grading standards, and demographic prediction; and (3) systematic evaluation of 24 state-of-the-art MLLMs. Our evaluations reveal both promise and limitations. Top-performing models achieved ~58% accuracy in disease screening under zero-shot settings, and performance remained suboptimal for challenging tasks like disease staging. We will publicly release the dataset, curation pipeline, and leaderboard to potentially advance ophthalmic AI applications and reduce the global burden of vision-threatening diseases.
title LMOD+: A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating Multimodal Large Language Models in Ophthalmology
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2509.25620