EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Sijing, Lin, Tianwei, Lin, Lingshuai, Zhang, Wenqiao, Liu, Jiang, Yang, Xiaoda, Li, Juncheng, He, Yucheng, Song, Xiaohui, Xiao, Jun, Zhuang, Yueting, Ooi, Beng Chin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916696232558592
author Li, Sijing
Lin, Tianwei
Lin, Lingshuai
Zhang, Wenqiao
Liu, Jiang
Yang, Xiaoda
Li, Juncheng
He, Yucheng
Song, Xiaohui
Xiao, Jun
Zhuang, Yueting
Ooi, Beng Chin
author_facet Li, Sijing
Lin, Tianwei
Lin, Lingshuai
Zhang, Wenqiao
Liu, Jiang
Yang, Xiaoda
Li, Juncheng
He, Yucheng
Song, Xiaohui
Xiao, Jun
Zhuang, Yueting
Ooi, Beng Chin
contents Medical Large Vision-Language Models (Med-LVLMs) demonstrate significant potential in healthcare, but their reliance on general medical data and coarse-grained global visual understanding limits them in intelligent ophthalmic diagnosis. Currently, intelligent ophthalmic diagnosis faces three major challenges: (i) Data. The lack of deeply annotated, high-quality, multi-modal ophthalmic visual instruction data; (ii) Benchmark. The absence of a comprehensive and systematic benchmark for evaluating diagnostic performance; (iii) Model. The difficulty of adapting holistic visual architectures to fine-grained, region-specific ophthalmic lesion identification. In this paper, we propose the Eyecare Kit, which systematically tackles the aforementioned three key challenges with the tailored dataset, benchmark and model: First, we construct a multi-agent data engine with real-life ophthalmology data to produce Eyecare-100K, a high-quality ophthalmic visual instruction dataset. Subsequently, we design Eyecare-Bench, a benchmark that comprehensively evaluates the overall performance of LVLMs on intelligent ophthalmic diagnosis tasks across multiple dimensions. Finally, we develop the EyecareGPT, optimized for fine-grained ophthalmic visual understanding thoroughly, which incorporates an adaptive resolution mechanism and a layer-wise dense connector. Extensive experimental results indicate that the EyecareGPT achieves state-of-the-art performance in a range of ophthalmic tasks, underscoring its significant potential for the advancement of open research in intelligent ophthalmic diagnosis. Our project is available at https://github.com/DCDmllm/EyecareGPT.
format Preprint
id arxiv_https___arxiv_org_abs_2504_13650
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
Li, Sijing
Lin, Tianwei
Lin, Lingshuai
Zhang, Wenqiao
Liu, Jiang
Yang, Xiaoda
Li, Juncheng
He, Yucheng
Song, Xiaohui
Xiao, Jun
Zhuang, Yueting
Ooi, Beng Chin
Computer Vision and Pattern Recognition
Medical Large Vision-Language Models (Med-LVLMs) demonstrate significant potential in healthcare, but their reliance on general medical data and coarse-grained global visual understanding limits them in intelligent ophthalmic diagnosis. Currently, intelligent ophthalmic diagnosis faces three major challenges: (i) Data. The lack of deeply annotated, high-quality, multi-modal ophthalmic visual instruction data; (ii) Benchmark. The absence of a comprehensive and systematic benchmark for evaluating diagnostic performance; (iii) Model. The difficulty of adapting holistic visual architectures to fine-grained, region-specific ophthalmic lesion identification. In this paper, we propose the Eyecare Kit, which systematically tackles the aforementioned three key challenges with the tailored dataset, benchmark and model: First, we construct a multi-agent data engine with real-life ophthalmology data to produce Eyecare-100K, a high-quality ophthalmic visual instruction dataset. Subsequently, we design Eyecare-Bench, a benchmark that comprehensively evaluates the overall performance of LVLMs on intelligent ophthalmic diagnosis tasks across multiple dimensions. Finally, we develop the EyecareGPT, optimized for fine-grained ophthalmic visual understanding thoroughly, which incorporates an adaptive resolution mechanism and a layer-wise dense connector. Extensive experimental results indicate that the EyecareGPT achieves state-of-the-art performance in a range of ophthalmic tasks, underscoring its significant potential for the advancement of open research in intelligent ophthalmic diagnosis. Our project is available at https://github.com/DCDmllm/EyecareGPT.
title EyecareGPT: Boosting Comprehensive Ophthalmology Understanding with Tailored Dataset, Benchmark and Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.13650