BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866909697987051520 |
|---|---|
| author | Srinivasan, Sahana Ai, Xuguang Lo, Thaddaeus Wai Soon Gilson, Aidan Zou, Minjie Zou, Ke Kim, Hyunjae Yang, Mingjia Pushpanathan, Krithi Yew, Samantha Loke, Wan Ting Goh, Jocelyn Chen, Yibing Kong, Yiming Fu, Emily Yuelei Hui, Michelle Ongyong Nwanyanwu, Kristen Dave, Amisha Li, Kelvin Zhenghao Sun, Chen-Hsin Chia, Mark Yang, Gabriel Dawei Wong, Wendy Meihua Chen, David Ziyou Liu, Dianbo Singer, Maxwell Antaki, Fares Del Priore, Lucian V Jonas, Jost Adelman, Ron Chen, Qingyu Tham, Yih-Chung |
| author_facet | Srinivasan, Sahana Ai, Xuguang Lo, Thaddaeus Wai Soon Gilson, Aidan Zou, Minjie Zou, Ke Kim, Hyunjae Yang, Mingjia Pushpanathan, Krithi Yew, Samantha Loke, Wan Ting Goh, Jocelyn Chen, Yibing Kong, Yiming Fu, Emily Yuelei Hui, Michelle Ongyong Nwanyanwu, Kristen Dave, Amisha Li, Kelvin Zhenghao Sun, Chen-Hsin Chia, Mark Yang, Gabriel Dawei Wong, Wendy Meihua Chen, David Ziyou Liu, Dianbo Singer, Maxwell Antaki, Fares Del Priore, Lucian V Jonas, Jost Adelman, Ron Chen, Qingyu Tham, Yih-Chung |
| contents | Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and comprehensive evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BELO assesses ophthalmology-related clinical accuracy and reasoning quality. Using keyword matching and a fine-tuned PubMedBERT model, we curated ophthalmology-specific multiple-choice-questions (MCQs) from diverse medical datasets (BCSC, MedMCQA, MedQA, BioASQ, and PubMedQA). The dataset underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by three senior ophthalmologists. To illustrate BELO's utility, we evaluated six LLMs (OpenAI o1, o3-mini, GPT-4o, DeepSeek-R1, Llama-3-8B, and Gemini 1.5 Pro) using accuracy, macro-F1, and five text-generation metrics (ROUGE-L, BERTScore, BARTScore, METEOR, and AlignScore). In a further evaluation involving human experts, two ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. BELO consists of 900 high-quality, expert-reviewed questions aggregated from five sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO dataset will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2507_15717 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning Srinivasan, Sahana Ai, Xuguang Lo, Thaddaeus Wai Soon Gilson, Aidan Zou, Minjie Zou, Ke Kim, Hyunjae Yang, Mingjia Pushpanathan, Krithi Yew, Samantha Loke, Wan Ting Goh, Jocelyn Chen, Yibing Kong, Yiming Fu, Emily Yuelei Hui, Michelle Ongyong Nwanyanwu, Kristen Dave, Amisha Li, Kelvin Zhenghao Sun, Chen-Hsin Chia, Mark Yang, Gabriel Dawei Wong, Wendy Meihua Chen, David Ziyou Liu, Dianbo Singer, Maxwell Antaki, Fares Del Priore, Lucian V Jonas, Jost Adelman, Ron Chen, Qingyu Tham, Yih-Chung Computation and Language Artificial Intelligence Current benchmarks evaluating large language models (LLMs) in ophthalmology are limited in scope and disproportionately prioritise accuracy. We introduce BELO (BEnchmarking LLMs for Ophthalmology), a standardized and comprehensive evaluation benchmark developed through multiple rounds of expert checking by 13 ophthalmologists. BELO assesses ophthalmology-related clinical accuracy and reasoning quality. Using keyword matching and a fine-tuned PubMedBERT model, we curated ophthalmology-specific multiple-choice-questions (MCQs) from diverse medical datasets (BCSC, MedMCQA, MedQA, BioASQ, and PubMedQA). The dataset underwent multiple rounds of expert checking. Duplicate and substandard questions were systematically removed. Ten ophthalmologists refined the explanations of each MCQ's correct answer. This was further adjudicated by three senior ophthalmologists. To illustrate BELO's utility, we evaluated six LLMs (OpenAI o1, o3-mini, GPT-4o, DeepSeek-R1, Llama-3-8B, and Gemini 1.5 Pro) using accuracy, macro-F1, and five text-generation metrics (ROUGE-L, BERTScore, BARTScore, METEOR, and AlignScore). In a further evaluation involving human experts, two ophthalmologists qualitatively reviewed 50 randomly selected outputs for accuracy, comprehensiveness, and completeness. BELO consists of 900 high-quality, expert-reviewed questions aggregated from five sources: BCSC (260), BioASQ (10), MedMCQA (572), MedQA (40), and PubMedQA (18). A public leaderboard has been established to promote transparent evaluation and reporting. Importantly, the BELO dataset will remain a hold-out, evaluation-only benchmark to ensure fair and reproducible comparisons of future models. |
| title | BEnchmarking LLMs for Ophthalmology (BELO) for Ophthalmological Knowledge and Reasoning |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2507.15717 |