RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Ning, Junzhi, Tang, Cheng, Zhou, Kaijing, Song, Diping, Liu, Lihao, Hu, Ming, Li, Wei, Xu, Huihui, Su, Yanzhou, Li, Tianbin, Liu, Jiyao, Ye, Jin, Zhang, Sheng, Ji, Yuanfeng, He, Junjun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915394548137984
author Ning, Junzhi
Tang, Cheng
Zhou, Kaijing
Song, Diping
Liu, Lihao
Hu, Ming
Li, Wei
Xu, Huihui
Su, Yanzhou
Li, Tianbin
Liu, Jiyao
Ye, Jin
Zhang, Sheng
Ji, Yuanfeng
He, Junjun
author_facet Ning, Junzhi
Tang, Cheng
Zhou, Kaijing
Song, Diping
Liu, Lihao
Hu, Ming
Li, Wei
Xu, Huihui
Su, Yanzhou
Li, Tianbin
Liu, Jiyao
Ye, Jin
Zhang, Sheng
Ji, Yuanfeng
He, Junjun
contents The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour Fundus Photographs (CFPs) largely rely on predefined disease labels, which restricts their ability to generate images that reflect fine-grained anatomical variations, subtle disease stages, and diverse pathological features beyond coarse class categories. To overcome these challenges, we first introduce an innovative pipeline that creates a large-scale, captioned retinal dataset comprising 1.4 million entries, called RetinaLogos-1400k. Specifically, RetinaLogos-1400k uses the visual language model(VLM) to describe retinal conditions and key structures, such as optic disc configuration, vascular distribution, nerve fibre layers, and pathological features. Building on this dataset, we employ a novel three-step training framework, RetinaLogos, which enables fine-grained semantic control over retinal images and accurately captures different stages of disease progression, subtle anatomical variations, and specific lesion types. Through extensive experiments, our method demonstrates superior performance across multiple datasets, with 62.07% of text-driven synthetic CFPs indistinguishable from real ones by ophthalmologists. Moreover, the synthetic data improves accuracy by 5%-10% in diabetic retinopathy grading and glaucoma detection. Codes are available at https://github.com/uni-medical/retina-text2cfp.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12887
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
Ning, Junzhi
Tang, Cheng
Zhou, Kaijing
Song, Diping
Liu, Lihao
Hu, Ming
Li, Wei
Xu, Huihui
Su, Yanzhou
Li, Tianbin
Liu, Jiyao
Ye, Jin
Zhang, Sheng
Ji, Yuanfeng
He, Junjun
Image and Video Processing
Computer Vision and Pattern Recognition
The scarcity of high-quality, labelled retinal imaging data, which presents a significant challenge in the development of machine learning models for ophthalmology, hinders progress in the field. Existing methods for synthesising Colour Fundus Photographs (CFPs) largely rely on predefined disease labels, which restricts their ability to generate images that reflect fine-grained anatomical variations, subtle disease stages, and diverse pathological features beyond coarse class categories. To overcome these challenges, we first introduce an innovative pipeline that creates a large-scale, captioned retinal dataset comprising 1.4 million entries, called RetinaLogos-1400k. Specifically, RetinaLogos-1400k uses the visual language model(VLM) to describe retinal conditions and key structures, such as optic disc configuration, vascular distribution, nerve fibre layers, and pathological features. Building on this dataset, we employ a novel three-step training framework, RetinaLogos, which enables fine-grained semantic control over retinal images and accurately captures different stages of disease progression, subtle anatomical variations, and specific lesion types. Through extensive experiments, our method demonstrates superior performance across multiple datasets, with 62.07% of text-driven synthetic CFPs indistinguishable from real ones by ophthalmologists. Moreover, the synthetic data improves accuracy by 5%-10% in diabetic retinopathy grading and glaucoma detection. Codes are available at https://github.com/uni-medical/retina-text2cfp.
title RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
topic Image and Video Processing
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2505.12887