Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866910797155794944 |
|---|---|
| author | Srinivasan, Sahana Ai, Xuguang Zou, Minjie Zou, Ke Kim, Hyunjae Lo, Thaddaeus Wai Soon Pushpanathan, Krithi Kong, Yiming Li, Anran Singer, Maxwell Jin, Kai Antaki, Fares Chen, David Ziyou Liu, Dianbo Adelman, Ron A. Chen, Qingyu Tham, Yih Chung |
| author_facet | Srinivasan, Sahana Ai, Xuguang Zou, Minjie Zou, Ke Kim, Hyunjae Lo, Thaddaeus Wai Soon Pushpanathan, Krithi Kong, Yiming Li, Anran Singer, Maxwell Jin, Kai Antaki, Fares Chen, David Ziyou Liu, Dianbo Adelman, Ron A. Chen, Qingyu Tham, Yih Chung |
| contents | Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions?
Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations.
Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_13949 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study Srinivasan, Sahana Ai, Xuguang Zou, Minjie Zou, Ke Kim, Hyunjae Lo, Thaddaeus Wai Soon Pushpanathan, Krithi Kong, Yiming Li, Anran Singer, Maxwell Jin, Kai Antaki, Fares Chen, David Ziyou Liu, Dianbo Adelman, Ron A. Chen, Qingyu Tham, Yih Chung Computation and Language Artificial Intelligence Question: What is the performance and reasoning ability of OpenAI o1 compared to other large language models in addressing ophthalmology-specific questions? Findings: This study evaluated OpenAI o1 and five LLMs using 6,990 ophthalmological questions from MedMCQA. O1 achieved the highest accuracy (0.88) and macro-F1 score but ranked third in reasoning capabilities based on text-generation metrics. Across subtopics, o1 ranked first in ``Lens'' and ``Glaucoma'' but second to GPT-4o in ``Corneal and External Diseases'', ``Vitreous and Retina'' and ``Oculoplastic and Orbital Diseases''. Subgroup analyses showed o1 performed better on queries with longer ground truth explanations. Meaning: O1's reasoning enhancements may not fully extend to ophthalmology, underscoring the need for domain-specific refinements to optimize performance in specialized fields like ophthalmology. |
| title | Can OpenAI o1 Reason Well in Ophthalmology? A 6,990-Question Head-to-Head Evaluation Study |
| topic | Computation and Language Artificial Intelligence |
| url | https://arxiv.org/abs/2501.13949 |