BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srinivasan, Sahana, Ai, Xuguang, Lo, Thaddaeus Wai Soon, Gilson, Aidan, Zou, Minjie, Zou, Ke, Kim, Hyunjae, Yang, Mingjia, Pushpanathan, Krithi, Yew, Samantha, Loke, Wan Ting, Goh, Jocelyn, Chen, Yibing, Kong, Yiming, Fu, Emily Yuelei, Hui, Michelle Ongyong, Nwanyanwu, Kristen, Dave, Amisha, Li, Kelvin Zhenghao, Sun, Chen-Hsin, Chia, Mark, Yang, Gabriel Dawei, Wong, Wendy Meihua, Chen, David Ziyou, Liu, Dianbo, Singer, Maxwell, Antaki, Fares, Del Priore, Lucian V, Jonas, Jost, Adelman, Ron, Chen, Qingyu, Tham, Yih-Chung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909697987051520
author Srinivasan, Sahana
Ai, Xuguang
Lo, Thaddaeus Wai Soon
Gilson, Aidan
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Yang, Mingjia
Pushpanathan, Krithi
Yew, Samantha
Loke, Wan Ting
Goh, Jocelyn
Chen, Yibing
Kong, Yiming
Fu, Emily Yuelei
Hui, Michelle Ongyong
Nwanyanwu, Kristen
Dave, Amisha
Li, Kelvin Zhenghao
Sun, Chen-Hsin
Chia, Mark
Yang, Gabriel Dawei
Wong, Wendy Meihua
Chen, David Ziyou
Liu, Dianbo
Singer, Maxwell
Antaki, Fares
Del Priore, Lucian V
Jonas, Jost
Adelman, Ron
Chen, Qingyu
Tham, Yih-Chung
author_facet Srinivasan, Sahana
Ai, Xuguang
Lo, Thaddaeus Wai Soon
Gilson, Aidan
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Yang, Mingjia
Pushpanathan, Krithi
Yew, Samantha
Loke, Wan Ting
Goh, Jocelyn
Chen, Yibing
Kong, Yiming
Fu, Emily Yuelei
Hui, Michelle Ongyong
Nwanyanwu, Kristen
Dave, Amisha
Li, Kelvin Zhenghao
Sun, Chen-Hsin
Chia, Mark
Yang, Gabriel Dawei
Wong, Wendy Meihua
Chen, David Ziyou
Liu, Dianbo
Singer, Maxwell
Antaki, Fares
Del Priore, Lucian V
Jonas, Jost
Adelman, Ron
Chen, Qingyu
Tham, Yih-Chung
contents Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and comprehensive evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BELO assesses ophthalmology-related clinical accuracy and reasoning quality. Using keyword matching and a fine-tuned PubMedBERT model, we curated ophthalmology-specific multiple-choice-questions (MCQs) from diverse medical datasets (BCSC, MedMCQA, MedQA, BioASQ, and PubMedQA). The dataset underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by three senior ophthalmologists. To illustrate BELO's utility, we evaluated six LLMs (OpenAI o1, o3-mini, GPT-4o, DeepSeek-R1, Llama-3-8B, and Gemini 1.5 Pro) using accuracy, macro-F1, and five text-generation metrics (ROUGE-L, BERTScore, BARTScore, METEOR, and AlignScore). In a further evaluation involving human experts, two ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. BELO consists of 900 high-quality, expert-reviewed questions aggregated from five sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO dataset will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning
Srinivasan, Sahana
Ai, Xuguang
Lo, Thaddaeus Wai Soon
Gilson, Aidan
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Yang, Mingjia
Pushpanathan, Krithi
Yew, Samantha
Loke, Wan Ting
Goh, Jocelyn
Chen, Yibing
Kong, Yiming
Fu, Emily Yuelei
Hui, Michelle Ongyong
Nwanyanwu, Kristen
Dave, Amisha
Li, Kelvin Zhenghao
Sun, Chen-Hsin
Chia, Mark
Yang, Gabriel Dawei
Wong, Wendy Meihua
Chen, David Ziyou
Liu, Dianbo
Singer, Maxwell
Antaki, Fares
Del Priore, Lucian V
Jonas, Jost
Adelman, Ron
Chen, Qingyu
Tham, Yih-Chung
Computation and Language
Artificial Intelligence
Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and comprehensive evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BELO assesses ophthalmology-related clinical accuracy and reasoning quality. Using keyword matching and a fine-tuned PubMedBERT model, we curated ophthalmology-specific multiple-choice-questions (MCQs) from diverse medical datasets (BCSC, MedMCQA, MedQA, BioASQ, and PubMedQA). The dataset underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by three senior ophthalmologists. To illustrate BELO's utility, we evaluated six LLMs (OpenAI o1, o3-mini, GPT-4o, DeepSeek-R1, Llama-3-8B, and Gemini 1.5 Pro) using accuracy, macro-F1, and five text-generation metrics (ROUGE-L, BERTScore, BARTScore, METEOR, and AlignScore). In a further evaluation involving human experts, two ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. BELO consists of 900 high-quality, expert-reviewed questions aggregated from five sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO dataset will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models.
title BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2507.15717