DualStreamFoveaNet: A Dual Stream Fusion Architecture with Anatomical Awareness for Robust Fovea Localization

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Song, Sifan, Wang, Jinfeng, Wang, Zilong, Wang, Hongxing, Su, Jionglong, Ding, Xiaowei, Dang, Kang
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914968609226752
author Song, Sifan
Wang, Jinfeng
Wang, Zilong
Wang, Hongxing
Su, Jionglong
Ding, Xiaowei
Dang, Kang
author_facet Song, Sifan
Wang, Jinfeng
Wang, Zilong
Wang, Hongxing
Su, Jionglong
Ding, Xiaowei
Dang, Kang
contents Accurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local anatomical landmarks around the fovea, the inability to robustly handle diseased retinal images, and the variations in image conditions. In this paper, we propose a novel transformer-based architecture called DualStreamFoveaNet (DSFN) for multi-cue fusion. This architecture explicitly incorporates long-range connections and global features using retina and vessel distributions for robust fovea localization. We introduce a spatial attention mechanism in the dual-stream encoder to extract and fuse self-learned anatomical information, focusing more on features distributed along blood vessels and significantly reducing computational costs by decreasing token numbers. Our extensive experiments show that the proposed architecture achieves state-of-the-art performance on two public datasets and one large-scale private dataset. Furthermore, we demonstrate that the DSFN is more robust on both normal and diseased retina images and has better generalization capacity in cross-dataset experiments.
format Preprint
id arxiv_https___arxiv_org_abs_2302_06961
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle DualStreamFoveaNet: A Dual Stream Fusion Architecture with Anatomical Awareness for Robust Fovea Localization
Song, Sifan
Wang, Jinfeng
Wang, Zilong
Wang, Hongxing
Su, Jionglong
Ding, Xiaowei
Dang, Kang
Computer Vision and Pattern Recognition
Artificial Intelligence
Accurate fovea localization is essential for analyzing retinal diseases to prevent irreversible vision loss. While current deep learning-based methods outperform traditional ones, they still face challenges such as the lack of local anatomical landmarks around the fovea, the inability to robustly handle diseased retinal images, and the variations in image conditions. In this paper, we propose a novel transformer-based architecture called DualStreamFoveaNet (DSFN) for multi-cue fusion. This architecture explicitly incorporates long-range connections and global features using retina and vessel distributions for robust fovea localization. We introduce a spatial attention mechanism in the dual-stream encoder to extract and fuse self-learned anatomical information, focusing more on features distributed along blood vessels and significantly reducing computational costs by decreasing token numbers. Our extensive experiments show that the proposed architecture achieves state-of-the-art performance on two public datasets and one large-scale private dataset. Furthermore, we demonstrate that the DSFN is more robust on both normal and diseased retina images and has better generalization capacity in cross-dataset experiments.
title DualStreamFoveaNet: A Dual Stream Fusion Architecture with Anatomical Awareness for Robust Fovea Localization
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2302.06961