Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tan, Ting Fang, Elangovan, Kabilan, Pollreisz, Andreas, Dy, Kevin Bryan, Ng, Wei Yan, Wong, Joy Le Yi, Liyuan, Jin, Ning, Chrystie Quek Wan, Hong, Ashley Shuen Ying, Thirunavukarasu, Arun James, Chang, Shelley Yin-His, Yao, Jie, Hong, Dylan, Zhaoran, Wang, Gupta, Amrita, Ting, Daniel SW
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910012702457856
author Tan, Ting Fang
Elangovan, Kabilan
Pollreisz, Andreas
Dy, Kevin Bryan
Ng, Wei Yan
Wong, Joy Le Yi
Liyuan, Jin
Ning, Chrystie Quek Wan
Hong, Ashley Shuen Ying
Thirunavukarasu, Arun James
Chang, Shelley Yin-His
Yao, Jie
Hong, Dylan
Zhaoran, Wang
Gupta, Amrita
Ting, Daniel SW
author_facet Tan, Ting Fang
Elangovan, Kabilan
Pollreisz, Andreas
Dy, Kevin Bryan
Ng, Wei Yan
Wong, Joy Le Yi
Liyuan, Jin
Ning, Chrystie Quek Wan
Hong, Ashley Shuen Ying
Thirunavukarasu, Arun James
Chang, Shelley Yin-His
Yao, Jie
Hong, Dylan
Zhaoran, Wang
Gupta, Amrita
Ting, Daniel SW
contents Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.
format Preprint
id arxiv_https___arxiv_org_abs_2602_05381
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
Tan, Ting Fang
Elangovan, Kabilan
Pollreisz, Andreas
Dy, Kevin Bryan
Ng, Wei Yan
Wong, Joy Le Yi
Liyuan, Jin
Ning, Chrystie Quek Wan
Hong, Ashley Shuen Ying
Thirunavukarasu, Arun James
Chang, Shelley Yin-His
Yao, Jie
Hong, Dylan
Zhaoran, Wang
Gupta, Amrita
Ting, Daniel SW
Artificial Intelligence
Domain specific large language models are increasingly used to support patient education, triage, and clinical decision making in ophthalmology, making rigorous evaluation essential to ensure safety and accuracy. This study evaluated four small medical LLMs Meerkat-7B, BioMistral-7B, OpenBioLLM-8B, and MedLLaMA3-v20 in answering ophthalmology related patient queries and assessed the feasibility of LLM based evaluation against clinician grading. In this cross sectional study, 180 ophthalmology patient queries were answered by each model, generating 2160 responses. Models were selected for parameter sizes under 10 billion to enable resource efficient deployment. Responses were evaluated by three ophthalmologists of differing seniority and by GPT-4-Turbo using the S.C.O.R.E. framework assessing safety, consensus and context, objectivity, reproducibility, and explainability, with ratings assigned on a five point Likert scale. Agreement between LLM and clinician grading was assessed using Spearman rank correlation, Kendall tau statistics, and kernel density estimate analyses. Meerkat-7B achieved the highest performance with mean scores of 3.44 from Senior Consultants, 4.08 from Consultants, and 4.18 from Residents. MedLLaMA3-v20 performed poorest, with 25.5 percent of responses containing hallucinations or clinically misleading content, including fabricated terminology. GPT-4-Turbo grading showed strong alignment with clinician assessments overall, with Spearman rho of 0.80 and Kendall tau of 0.67, though Senior Consultants graded more conservatively. Overall, medical LLMs demonstrated potential for safe ophthalmic question answering, but gaps remained in clinical depth and consensus, supporting the feasibility of LLM based evaluation for large scale benchmarking and the need for hybrid automated and clinician review frameworks to guide safe clinical deployment.
title Clinical Validation of Medical-based Large Language Model Chatbots on Ophthalmic Patient Queries with LLM-based Evaluation
topic Artificial Intelligence
url https://arxiv.org/abs/2602.05381