Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Chenjun, Wan, Cheng, Lux, Laurin, Berger, Alexander, Rosen, Richard B., Menten, Martin J., Paetzold, Johannes C.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918245470044160
author Li, Chenjun
Wan, Cheng
Lux, Laurin
Berger, Alexander
Rosen, Richard B.
Menten, Martin J.
Paetzold, Johannes C.
author_facet Li, Chenjun
Wan, Cheng
Lux, Laurin
Berger, Alexander
Rosen, Richard B.
Menten, Martin J.
Paetzold, Johannes C.
contents Vision-Language Models (VLMs) offer a promising path toward interpretable medical diagnosis by allowing users to ask about clinical explanations alongside predictions and across different modalities. However, training VLMs for detailed reasoning requires large-scale image-text datasets. In many specialized domains, for example in reading Optical Coherence Tomography Angiography (OCTA) images, such precise text with grounded description of pathologies is scarce or even non-existent. To overcome this bottleneck, we introduce Synthetic Vasculature Reasoning (SVR), a framework that controllably synthesizes images and corresponding text, specifically: realistic retinal vasculature with Diabetic Retinopathy (DR) features: capillary dropout, microaneurysms, neovascularization, and tortuosity, while automatically generating granular reasoning texts. Based on this we curate OCTA-100K-SVR, an OCTA image-reasoning dataset with 100,000 pairs. Our experiments show that a general-purpose VLM (Qwen3-VL-8b) trained on the dataset achieves a zero-shot balanced classification accuracy of 89.67% on real OCTA images, outperforming supervised baselines. Through human expert evaluation we also demonstrate that it significantly enhances explanation quality and pathology localization on clinical data.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11060
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
Li, Chenjun
Wan, Cheng
Lux, Laurin
Berger, Alexander
Rosen, Richard B.
Menten, Martin J.
Paetzold, Johannes C.
Computer Vision and Pattern Recognition
Vision-Language Models (VLMs) offer a promising path toward interpretable medical diagnosis by allowing users to ask about clinical explanations alongside predictions and across different modalities. However, training VLMs for detailed reasoning requires large-scale image-text datasets. In many specialized domains, for example in reading Optical Coherence Tomography Angiography (OCTA) images, such precise text with grounded description of pathologies is scarce or even non-existent. To overcome this bottleneck, we introduce Synthetic Vasculature Reasoning (SVR), a framework that controllably synthesizes images and corresponding text, specifically: realistic retinal vasculature with Diabetic Retinopathy (DR) features: capillary dropout, microaneurysms, neovascularization, and tortuosity, while automatically generating granular reasoning texts. Based on this we curate OCTA-100K-SVR, an OCTA image-reasoning dataset with 100,000 pairs. Our experiments show that a general-purpose VLM (Qwen3-VL-8b) trained on the dataset achieves a zero-shot balanced classification accuracy of 89.67% on real OCTA images, outperforming supervised baselines. Through human expert evaluation we also demonstrate that it significantly enhances explanation quality and pathology localization on clinical data.
title Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11060