Confidence-aware multi-modality learning for eye disease screening

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zou, Ke, Lin, Tian, Han, Zongbo, Wang, Meng, Yuan, Xuedong, Chen, Haoyu, Zhang, Changqing, Shen, Xiaojing, Fu, Huazhu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913366994321408
author Zou, Ke
Lin, Tian
Han, Zongbo
Wang, Meng
Yuan, Xuedong
Chen, Haoyu
Zhang, Changqing
Shen, Xiaojing
Fu, Huazhu
author_facet Zou, Ke
Lin, Tian
Han, Zongbo
Wang, Meng
Yuan, Xuedong
Chen, Haoyu
Zhang, Changqing
Shen, Xiaojing
Fu, Huazhu
contents Multi-modal ophthalmic image classification plays a key role in diagnosing eye diseases, as it integrates information from different sources to complement their respective performances. However, recent improvements have mainly focused on accuracy, often neglecting the importance of confidence and robustness in predictions for diverse modalities. In this study, we propose a novel multi-modality evidential fusion pipeline for eye disease screening. It provides a measure of confidence for each modality and elegantly integrates the multi-modality information using a multi-distribution fusion perspective. Specifically, our method first utilizes normal inverse gamma prior distributions over pre-trained models to learn both aleatoric and epistemic uncertainty for uni-modality. Then, the normal inverse gamma distribution is analyzed as the Student's t distribution. Furthermore, within a confidence-aware fusion framework, we propose a mixture of Student's t distributions to effectively integrate different modalities, imparting the model with heavy-tailed properties and enhancing its robustness and reliability. More importantly, the confidence-aware multi-modality ranking regularization term induces the model to more reasonably rank the noisy single-modal and fused-modal confidence, leading to improved reliability and accuracy. Experimental results on both public and internal datasets demonstrate that our model excels in robustness, particularly in challenging scenarios involving Gaussian noise and modality missing conditions. Moreover, our model exhibits strong generalization capabilities to out-of-distribution data, underscoring its potential as a promising solution for multimodal eye disease screening.
format Preprint
id arxiv_https___arxiv_org_abs_2405_18167
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Confidence-aware multi-modality learning for eye disease screening
Zou, Ke
Lin, Tian
Han, Zongbo
Wang, Meng
Yuan, Xuedong
Chen, Haoyu
Zhang, Changqing
Shen, Xiaojing
Fu, Huazhu
Image and Video Processing
Computer Vision and Pattern Recognition
Multi-modal ophthalmic image classification plays a key role in diagnosing eye diseases, as it integrates information from different sources to complement their respective performances. However, recent improvements have mainly focused on accuracy, often neglecting the importance of confidence and robustness in predictions for diverse modalities. In this study, we propose a novel multi-modality evidential fusion pipeline for eye disease screening. It provides a measure of confidence for each modality and elegantly integrates the multi-modality information using a multi-distribution fusion perspective. Specifically, our method first utilizes normal inverse gamma prior distributions over pre-trained models to learn both aleatoric and epistemic uncertainty for uni-modality. Then, the normal inverse gamma distribution is analyzed as the Student's t distribution. Furthermore, within a confidence-aware fusion framework, we propose a mixture of Student's t distributions to effectively integrate different modalities, imparting the model with heavy-tailed properties and enhancing its robustness and reliability. More importantly, the confidence-aware multi-modality ranking regularization term induces the model to more reasonably rank the noisy single-modal and fused-modal confidence, leading to improved reliability and accuracy. Experimental results on both public and internal datasets demonstrate that our model excels in robustness, particularly in challenging scenarios involving Gaussian noise and modality missing conditions. Moreover, our model exhibits strong generalization capabilities to out-of-distribution data, underscoring its potential as a promising solution for multimodal eye disease screening.
title Confidence-aware multi-modality learning for eye disease screening
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2405.18167