MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Ruiqi, Zhang, Chenran, Zhang, Jianle, Zhou, Yi, Zhou, Tao, Fu, Huazhu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909752450088960
author Wu, Ruiqi
Zhang, Chenran
Zhang, Jianle
Zhou, Yi
Zhou, Tao
Fu, Huazhu
author_facet Wu, Ruiqi
Zhang, Chenran
Zhang, Jianle
Zhou, Yi
Zhou, Tao
Fu, Huazhu
contents Current fundus image analysis models are predominantly built for specific tasks relying on individual datasets. The learning process is usually based on data-driven paradigm without prior knowledge, resulting in poor transferability and generalizability. To address this issue, we propose MM-Retinal, a multi-modal dataset that encompasses high-quality image-text pairs collected from professional fundus diagram books. Moreover, enabled by MM-Retinal, we present a novel Knowledge-enhanced foundational pretraining model which incorporates Fundus Image-Text expertise, called KeepFIT. It is designed with image similarity-guided text revision and mixed training strategy to infuse expert knowledge. Our proposed fundus foundation model achieves state-of-the-art performance across six unseen downstream tasks and holds excellent generalization ability in zero-shot and few-shot scenarios. MM-Retinal and KeepFIT are available at https://github.com/lxirich/MM-Retinal.
format Preprint
id arxiv_https___arxiv_org_abs_2405_11793
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
Wu, Ruiqi
Zhang, Chenran
Zhang, Jianle
Zhou, Yi
Zhou, Tao
Fu, Huazhu
Computer Vision and Pattern Recognition
Current fundus image analysis models are predominantly built for specific tasks relying on individual datasets. The learning process is usually based on data-driven paradigm without prior knowledge, resulting in poor transferability and generalizability. To address this issue, we propose MM-Retinal, a multi-modal dataset that encompasses high-quality image-text pairs collected from professional fundus diagram books. Moreover, enabled by MM-Retinal, we present a novel Knowledge-enhanced foundational pretraining model which incorporates Fundus Image-Text expertise, called KeepFIT. It is designed with image similarity-guided text revision and mixed training strategy to infuse expert knowledge. Our proposed fundus foundation model achieves state-of-the-art performance across six unseen downstream tasks and holds excellent generalization ability in zero-shot and few-shot scenarios. MM-Retinal and KeepFIT are available at https://github.com/lxirich/MM-Retinal.
title MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.11793