Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Yukun, Nderitu, Paul, Goh, Jocelyn Hui Lin, Engelmann, Justin, Wagner, Siegfried K., Ran, Anran, Jiang, Hongyang, Ju, Lie, Zou, Ke, Srinivasan, Sahana, Kim, Hyunmin, Ninomiya, Takahiro, Wang, Zheyuan, Yang, Gabriel Dawei, Ruffell, Eden, Williamson, Dominic, Santos, Rui, Somfai, Gabor Mark, Cheung, Carol Y., Wong, Tien Yin, Alexander, Daniel C., Tham, Yih Chung, Keane, Pearse A.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915477909929984
author Zhou, Yukun
Nderitu, Paul
Goh, Jocelyn Hui Lin
Engelmann, Justin
Wagner, Siegfried K.
Ran, Anran
Jiang, Hongyang
Ju, Lie
Zou, Ke
Srinivasan, Sahana
Kim, Hyunmin
Ninomiya, Takahiro
Wang, Zheyuan
Yang, Gabriel Dawei
Ruffell, Eden
Williamson, Dominic
Santos, Rui
Somfai, Gabor Mark
Cheung, Carol Y.
Wong, Tien Yin
Alexander, Daniel C.
Tham, Yih Chung
Keane, Pearse A.
author_facet Zhou, Yukun
Nderitu, Paul
Goh, Jocelyn Hui Lin
Engelmann, Justin
Wagner, Siegfried K.
Ran, Anran
Jiang, Hongyang
Ju, Lie
Zou, Ke
Srinivasan, Sahana
Kim, Hyunmin
Ninomiya, Takahiro
Wang, Zheyuan
Yang, Gabriel Dawei
Ruffell, Eden
Williamson, Dominic
Santos, Rui
Somfai, Gabor Mark
Cheung, Carol Y.
Wong, Tien Yin
Alexander, Daniel C.
Tham, Yih Chung
Keane, Pearse A.
contents Medical foundation models, pre-trained with large-scale clinical data, demonstrate strong performance in diverse clinically relevant applications. RETFound, trained on nearly one million retinal images, exemplifies this approach in applications with retinal images. However, the emergence of increasingly powerful and multifold larger generalist foundation models such as DINOv2 and DINOv3 raises the question of whether domain-specific pre-training remains essential, and if so, what gap persists. To investigate this, we systematically evaluated the adaptability of DINOv2 and DINOv3 in retinal image applications, compared to two specialist RETFound models, RETFound-MAE and RETFound-DINOv2. We assessed performance on ocular disease detection and systemic disease prediction using two adaptation strategies: fine-tuning and linear probing. Data efficiency and adaptation efficiency were further analysed to characterise trade-offs between predictive performance and computational cost. Our results show that although scaling generalist models yields strong adaptability across diverse tasks, RETFound-DINOv2 consistently outperforms these generalist foundation models in ocular-disease detection and oculomics tasks, demonstrating stronger generalisability and data efficiency. These findings suggest that specialist retinal foundation models remain the most effective choice for clinical applications, while the narrowing gap with generalist foundation models suggests that continued data and model scaling can deliver domain-relevant gains and position them as strong foundations for future medical foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2509_03421
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
Zhou, Yukun
Nderitu, Paul
Goh, Jocelyn Hui Lin
Engelmann, Justin
Wagner, Siegfried K.
Ran, Anran
Jiang, Hongyang
Ju, Lie
Zou, Ke
Srinivasan, Sahana
Kim, Hyunmin
Ninomiya, Takahiro
Wang, Zheyuan
Yang, Gabriel Dawei
Ruffell, Eden
Williamson, Dominic
Santos, Rui
Somfai, Gabor Mark
Cheung, Carol Y.
Wong, Tien Yin
Alexander, Daniel C.
Tham, Yih Chung
Keane, Pearse A.
Image and Video Processing
Computer Vision and Pattern Recognition
J.3; I.2.10
Medical foundation models, pre-trained with large-scale clinical data, demonstrate strong performance in diverse clinically relevant applications. RETFound, trained on nearly one million retinal images, exemplifies this approach in applications with retinal images. However, the emergence of increasingly powerful and multifold larger generalist foundation models such as DINOv2 and DINOv3 raises the question of whether domain-specific pre-training remains essential, and if so, what gap persists. To investigate this, we systematically evaluated the adaptability of DINOv2 and DINOv3 in retinal image applications, compared to two specialist RETFound models, RETFound-MAE and RETFound-DINOv2. We assessed performance on ocular disease detection and systemic disease prediction using two adaptation strategies: fine-tuning and linear probing. Data efficiency and adaptation efficiency were further analysed to characterise trade-offs between predictive performance and computational cost. Our results show that although scaling generalist models yields strong adaptability across diverse tasks, RETFound-DINOv2 consistently outperforms these generalist foundation models in ocular-disease detection and oculomics tasks, demonstrating stronger generalisability and data efficiency. These findings suggest that specialist retinal foundation models remain the most effective choice for clinical applications, while the narrowing gap with generalist foundation models suggests that continued data and model scaling can deliver domain-relevant gains and position them as strong foundations for future medical foundation models.
title Generalist versus Specialist Vision Foundation Models for Ocular Disease and Oculomics
topic Image and Video Processing
Computer Vision and Pattern Recognition
J.3; I.2.10
url https://arxiv.org/abs/2509.03421