Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Fan, Hou, Junlin, Zhao, Ruiwei, Feng, Rui, Zou, Haidong, Lu, Lina, Xu, Yi, Zhang, Juzhao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909374679613440
author Xiao, Fan
Hou, Junlin
Zhao, Ruiwei
Feng, Rui
Zou, Haidong
Lu, Lina
Xu, Yi
Zhang, Juzhao
author_facet Xiao, Fan
Hou, Junlin
Zhao, Ruiwei
Feng, Rui
Zou, Haidong
Lu, Lina
Xu, Yi
Zhang, Juzhao
contents Diabetic retinopathy (DR) is a leading cause of blindness worldwide and a common complication of diabetes. As two different imaging tools for DR grading, color fundus photography (CFP) and infrared fundus photography (IFP) are highly-correlated and complementary in clinical applications. To the best of our knowledge, this is the first study that explores a novel multi-modal deep learning framework to fuse the information from CFP and IFP towards more accurate DR grading. Specifically, we construct a dual-stream architecture Cross-Fundus Transformer (CFT) to fuse the ViT-based features of two fundus image modalities. In particular, a meticulously engineered Cross-Fundus Attention (CFA) module is introduced to capture the correspondence between CFP and IFP images. Moreover, we adopt both the single-modality and multi-modality supervisions to maximize the overall performance for DR grading. Extensive experiments on a clinical dataset consisting of 1,713 pairs of multi-modal fundus images demonstrate the superiority of our proposed method. Our code will be released for public access.
format Preprint
id arxiv_https___arxiv_org_abs_2411_00726
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract
Xiao, Fan
Hou, Junlin
Zhao, Ruiwei
Feng, Rui
Zou, Haidong
Lu, Lina
Xu, Yi
Zhang, Juzhao
Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Diabetic retinopathy (DR) is a leading cause of blindness worldwide and a common complication of diabetes. As two different imaging tools for DR grading, color fundus photography (CFP) and infrared fundus photography (IFP) are highly-correlated and complementary in clinical applications. To the best of our knowledge, this is the first study that explores a novel multi-modal deep learning framework to fuse the information from CFP and IFP towards more accurate DR grading. Specifically, we construct a dual-stream architecture Cross-Fundus Transformer (CFT) to fuse the ViT-based features of two fundus image modalities. In particular, a meticulously engineered Cross-Fundus Attention (CFA) module is introduced to capture the correspondence between CFP and IFP images. Moreover, we adopt both the single-modality and multi-modality supervisions to maximize the overall performance for DR grading. Extensive experiments on a clinical dataset consisting of 1,713 pairs of multi-modal fundus images demonstrate the superiority of our proposed method. Our code will be released for public access.
title Cross-Fundus Transformer for Multi-modal Diabetic Retinopathy Grading with Cataract
topic Image and Video Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.00726