Performance of Large Language Models in Answering Critical Care Medicine Questions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alwakeel, Mahmoud, Nagori, Aditya, Wong, An-Kwok Ian, Chaisson, Neal, Krishnamoorthy, Vijay, Kamaleswaran, Rishikesan
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911172268130304
author Alwakeel, Mahmoud
Nagori, Aditya
Wong, An-Kwok Ian
Chaisson, Neal
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
author_facet Alwakeel, Mahmoud
Nagori, Aditya
Wong, An-Kwok Ian
Chaisson, Neal
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
contents Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.
format Preprint
id arxiv_https___arxiv_org_abs_2509_19344
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance of Large Language Models in Answering Critical Care Medicine Questions
Alwakeel, Mahmoud
Nagori, Aditya
Wong, An-Kwok Ian
Chaisson, Neal
Krishnamoorthy, Vijay
Kamaleswaran, Rishikesan
Computation and Language
Large Language Models have been tested on medical student-level questions, but their performance in specialized fields like Critical Care Medicine (CCM) is less explored. This study evaluated Meta-Llama 3.1 models (8B and 70B parameters) on 871 CCM questions. Llama3.1:70B outperformed 8B by 30%, with 60% average accuracy. Performance varied across domains, highest in Research (68.4%) and lowest in Renal (47.9%), highlighting the need for broader future work to improve models across various subspecialty domains.
title Performance of Large Language Models in Answering Critical Care Medicine Questions
topic Computation and Language
url https://arxiv.org/abs/2509.19344