Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Srinivasan, Sahana, Ai, Xuguang, Zou, Minjie, Zou, Ke, Kim, Hyunjae, Lo, Thaddaeus Wai Soon, Pushpanathan, Krithi, Kong, Yiming, Li, Anran, Singer, Maxwell, Jin, Kai, Antaki, Fares, Chen, David Ziyou, Liu, Dianbo, Adelman, Ron A., Chen, Qingyu, Tham, Yih Chung
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910797155794944
author Srinivasan, Sahana
Ai, Xuguang
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Lo, Thaddaeus Wai Soon
Pushpanathan, Krithi
Kong, Yiming
Li, Anran
Singer, Maxwell
Jin, Kai
Antaki, Fares
Chen, David Ziyou
Liu, Dianbo
Adelman, Ron A.
Chen, Qingyu
Tham, Yih Chung
author_facet Srinivasan, Sahana
Ai, Xuguang
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Lo, Thaddaeus Wai Soon
Pushpanathan, Krithi
Kong, Yiming
Li, Anran
Singer, Maxwell
Jin, Kai
Antaki, Fares
Chen, David Ziyou
Liu, Dianbo
Adelman, Ron A.
Chen, Qingyu
Tham, Yih Chung
contents Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions? Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations. Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology.
format Preprint
id arxiv_https___arxiv_org_abs_2501_13949
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
Srinivasan, Sahana
Ai, Xuguang
Zou, Minjie
Zou, Ke
Kim, Hyunjae
Lo, Thaddaeus Wai Soon
Pushpanathan, Krithi
Kong, Yiming
Li, Anran
Singer, Maxwell
Jin, Kai
Antaki, Fares
Chen, David Ziyou
Liu, Dianbo
Adelman, Ron A.
Chen, Qingyu
Tham, Yih Chung
Computation and Language
Artificial Intelligence
Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions? Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations. Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology.
title Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2501.13949