nnMobileNet: Rethinking CNN for Retinopathy Research

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Wenhui, Qiu, Peijie, Chen, Xiwen, Li, Xin, Lepore, Natasha, Dumitrascu, Oana M., Wang, Yalin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913316277846016
author Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Li, Xin
Lepore, Natasha
Dumitrascu, Oana M.
Wang, Yalin
author_facet Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Li, Xin
Lepore, Natasha
Dumitrascu, Oana M.
Wang, Yalin
contents Over the past few decades, convolutional neural networks (CNNs) have been at the forefront of the detection and tracking of various retinal diseases (RD). Despite their success, the emergence of vision transformers (ViT) in the 2020s has shifted the trajectory of RD model development. The leading-edge performance of ViT-based models in RD can be largely credited to their scalability-their ability to improve as more parameters are added. As a result, ViT-based models tend to outshine traditional CNNs in RD applications, albeit at the cost of increased data and computational demands. ViTs also differ from CNNs in their approach to processing images, working with patches rather than local regions, which can complicate the precise localization of small, variably presented lesions in RD. In our study, we revisited and updated the architecture of a CNN model, specifically MobileNet, to enhance its utility in RD diagnostics. We found that an optimized MobileNet, through selective modifications, can surpass ViT-based models in various RD benchmarks, including diabetic retinopathy grading, detection of multiple fundus diseases, and classification of diabetic macular edema. The code is available at https://github.com/Retinal-Research/NN-MOBILENET
format Preprint
id arxiv_https___arxiv_org_abs_2306_01289
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle nnMobileNet: Rethinking CNN for Retinopathy Research
Zhu, Wenhui
Qiu, Peijie
Chen, Xiwen
Li, Xin
Lepore, Natasha
Dumitrascu, Oana M.
Wang, Yalin
Image and Video Processing
Computer Vision and Pattern Recognition
Over the past few decades, convolutional neural networks (CNNs) have been at the forefront of the detection and tracking of various retinal diseases (RD). Despite their success, the emergence of vision transformers (ViT) in the 2020s has shifted the trajectory of RD model development. The leading-edge performance of ViT-based models in RD can be largely credited to their scalability-their ability to improve as more parameters are added. As a result, ViT-based models tend to outshine traditional CNNs in RD applications, albeit at the cost of increased data and computational demands. ViTs also differ from CNNs in their approach to processing images, working with patches rather than local regions, which can complicate the precise localization of small, variably presented lesions in RD. In our study, we revisited and updated the architecture of a CNN model, specifically MobileNet, to enhance its utility in RD diagnostics. We found that an optimized MobileNet, through selective modifications, can surpass ViT-based models in various RD benchmarks, including diabetic retinopathy grading, detection of multiple fundus diseases, and classification of diabetic macular edema. The code is available at https://github.com/Retinal-Research/NN-MOBILENET
title nnMobileNet: Rethinking CNN for Retinopathy Research
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2306.01289