GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Cheng, Xie, Weizheng, Kooner, Karanjit, Lee, Tsengdar, Wang, Jui-Kai, Zhang, Jia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909723967619072
author Huang, Cheng
Xie, Weizheng
Kooner, Karanjit
Lee, Tsengdar
Wang, Jui-Kai
Zhang, Jia
author_facet Huang, Cheng
Xie, Weizheng
Kooner, Karanjit
Lee, Tsengdar
Wang, Jui-Kai
Zhang, Jia
contents Early and accurate detection of glaucoma is critical to prevent irreversible vision loss. However, existing methods often rely on unimodal data and lack interpretability, limiting their clinical utility. In this paper, we present GlaBoost, a multimodal gradient boosting framework that integrates structured clinical features, fundus image embeddings, and expert-curated textual descriptions for glaucoma risk prediction. GlaBoost extracts high-level visual representations from retinal fundus photographs using a pretrained convolutional encoder and encodes free-text neuroretinal rim assessments using a transformer-based language model. These heterogeneous signals, combined with manually assessed risk scores and quantitative ophthalmic indicators, are fused into a unified feature space for classification via an enhanced XGBoost model. Experiments conducted on a real-world annotated dataset demonstrate that GlaBoost significantly outperforms baseline models, achieving a validation accuracy of 98.71%. Feature importance analysis reveals clinically consistent patterns, with cup-to-disc ratio, rim pallor, and specific textual embeddings contributing most to model decisions. GlaBoost offers a transparent and scalable solution for interpretable glaucoma diagnosis and can be extended to other ophthalmic disorders.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03750
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
Huang, Cheng
Xie, Weizheng
Kooner, Karanjit
Lee, Tsengdar
Wang, Jui-Kai
Zhang, Jia
Machine Learning
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
Image and Video Processing
Early and accurate detection of glaucoma is critical to prevent irreversible vision loss. However, existing methods often rely on unimodal data and lack interpretability, limiting their clinical utility. In this paper, we present GlaBoost, a multimodal gradient boosting framework that integrates structured clinical features, fundus image embeddings, and expert-curated textual descriptions for glaucoma risk prediction. GlaBoost extracts high-level visual representations from retinal fundus photographs using a pretrained convolutional encoder and encodes free-text neuroretinal rim assessments using a transformer-based language model. These heterogeneous signals, combined with manually assessed risk scores and quantitative ophthalmic indicators, are fused into a unified feature space for classification via an enhanced XGBoost model. Experiments conducted on a real-world annotated dataset demonstrate that GlaBoost significantly outperforms baseline models, achieving a validation accuracy of 98.71%. Feature importance analysis reveals clinically consistent patterns, with cup-to-disc ratio, rim pallor, and specific textual embeddings contributing most to model decisions. GlaBoost offers a transparent and scalable solution for interpretable glaucoma diagnosis and can be extended to other ophthalmic disorders.
title GlaBoost: A multimodal Structured Framework for Glaucoma Risk Stratification
topic Machine Learning
Computational Engineering, Finance, and Science
Computer Vision and Pattern Recognition
Image and Video Processing
url https://arxiv.org/abs/2508.03750