Multi-Scale Target-Aware Representation Learning for Fundus Image Enhancement

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Wu, Haofan, Huang, Yin, Wu, Yuqing, Yang, Qiuyu, Wang, Bingfang, Zhang, Li, Khan, Muhammad Fahadullah, Zia, Ali, Memon, M. Saleh, Bukhari, Syed Sohail, Memon, Abdul Fattah, Ji, Daizong, Zhang, Ya, Mustafa, Ghulam, Fang, Yin
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915619554721792
author Wu, Haofan
Huang, Yin
Wu, Yuqing
Yang, Qiuyu
Wang, Bingfang
Zhang, Li
Khan, Muhammad Fahadullah
Zia, Ali
Memon, M. Saleh
Bukhari, Syed Sohail
Memon, Abdul Fattah
Ji, Daizong
Zhang, Ya
Mustafa, Ghulam
Fang, Yin
author_facet Wu, Haofan
Huang, Yin
Wu, Yuqing
Yang, Qiuyu
Wang, Bingfang
Zhang, Li
Khan, Muhammad Fahadullah
Zia, Ali
Memon, M. Saleh
Bukhari, Syed Sohail
Memon, Abdul Fattah
Ji, Daizong
Zhang, Ya
Mustafa, Ghulam
Fang, Yin
contents High-quality fundus images provide essential anatomical information for clinical screening and ophthalmic disease diagnosis. Yet, due to hardware limitations, operational variability, and patient compliance, fundus images often suffer from low resolution and signal-to-noise ratio. Recent years have witnessed promising progress in fundus image enhancement. However, existing works usually focus on restoring structural details or global characteristics of fundus images, lacking a unified image enhancement framework to recover comprehensive multi-scale information. Moreover, few methods pinpoint the target of image enhancement, e.g., lesions, which is crucial for medical image-based diagnosis. To address these challenges, we propose a multi-scale target-aware representation learning framework (MTRL-FIE) for efficient fundus image enhancement. Specifically, we propose a multi-scale feature encoder (MFE) that employs wavelet decomposition to embed both low-frequency structural information and high-frequency details. Next, we design a structure-preserving hierarchical decoder (SHD) to fuse multi-scale feature embeddings for real fundus image restoration. SHD integrates hierarchical fusion and group attention mechanisms to achieve adaptive feature fusion while retaining local structural smoothness. Meanwhile, a target-aware feature aggregation (TFA) module is used to enhance pathological regions and reduce artifacts. Experimental results on multiple fundus image datasets demonstrate the effectiveness and generalizability of MTRL-FIE for fundus image enhancement. Compared to state-of-the-art methods, MTRL-FIE achieves superior enhancement performance with a more lightweight architecture. Furthermore, our approach generalizes to other ophthalmic image processing tasks without supervised fine-tuning, highlighting its potential for clinical applications.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01831
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Multi-Scale Target-Aware Representation Learning for Fundus Image Enhancement
Wu, Haofan
Huang, Yin
Wu, Yuqing
Yang, Qiuyu
Wang, Bingfang
Zhang, Li
Khan, Muhammad Fahadullah
Zia, Ali
Memon, M. Saleh
Bukhari, Syed Sohail
Memon, Abdul Fattah
Ji, Daizong
Zhang, Ya
Mustafa, Ghulam
Fang, Yin
Image and Video Processing
Computer Vision and Pattern Recognition
High-quality fundus images provide essential anatomical information for clinical screening and ophthalmic disease diagnosis. Yet, due to hardware limitations, operational variability, and patient compliance, fundus images often suffer from low resolution and signal-to-noise ratio. Recent years have witnessed promising progress in fundus image enhancement. However, existing works usually focus on restoring structural details or global characteristics of fundus images, lacking a unified image enhancement framework to recover comprehensive multi-scale information. Moreover, few methods pinpoint the target of image enhancement, e.g., lesions, which is crucial for medical image-based diagnosis. To address these challenges, we propose a multi-scale target-aware representation learning framework (MTRL-FIE) for efficient fundus image enhancement. Specifically, we propose a multi-scale feature encoder (MFE) that employs wavelet decomposition to embed both low-frequency structural information and high-frequency details. Next, we design a structure-preserving hierarchical decoder (SHD) to fuse multi-scale feature embeddings for real fundus image restoration. SHD integrates hierarchical fusion and group attention mechanisms to achieve adaptive feature fusion while retaining local structural smoothness. Meanwhile, a target-aware feature aggregation (TFA) module is used to enhance pathological regions and reduce artifacts. Experimental results on multiple fundus image datasets demonstrate the effectiveness and generalizability of MTRL-FIE for fundus image enhancement. Compared to state-of-the-art methods, MTRL-FIE achieves superior enhancement performance with a more lightweight architecture. Furthermore, our approach generalizes to other ophthalmic image processing tasks without supervised fine-tuning, highlighting its potential for clinical applications.
title Multi-Scale Target-Aware Representation Learning for Fundus Image Enhancement
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.01831