Can LLMs Score Medical Diagnoses and Clinical Reasoning as well as Expert Panels?

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rouillard, Amy, Mundia, Sitwala, Camara, Linda, Gramanie, Michael Cameron, Dangor, Ziyaad, Kalla, Ismail, Madhi, Shabir A., Morar, Kajal, Ncube, Marlvin T., Saloojee, Haroon, Bassett, Bruce A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910138080690176
author Rouillard, Amy
Mundia, Sitwala
Camara, Linda
Gramanie, Michael Cameron
Dangor, Ziyaad
Kalla, Ismail
Madhi, Shabir A.
Morar, Kajal
Ncube, Marlvin T.
Saloojee, Haroon
Bassett, Bruce A.
author_facet Rouillard, Amy
Mundia, Sitwala
Camara, Linda
Gramanie, Michael Cameron
Dangor, Ziyaad
Kalla, Ismail
Madhi, Shabir A.
Morar, Kajal
Ncube, Marlvin T.
Saloojee, Haroon
Bassett, Bruce A.
contents Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM jury composed of three frontier AI models scoring 3333 diagnoses on 300 real-world middle-income country (MIC) hospital cases. Model performance was benchmarked against expert clinician panel and independent human re-scoring panel evaluations. Both LLM and clinician-generated diagnoses are scored across four dimensions: diagnosis, differential diagnosis, clinical reasoning and negative treatment risk. For each of these, we assess scoring difference, inter-rater agreement, scoring stability, severe safety errors and the effect of post-hoc calibration. We find that: (i) the uncalibrated LLM jury scores are systematically lower than clinician panels scores; (ii) the LLM Jury preserves ordinal agreement and exhibits better concordance with the primary expert panels than the human expert re-score panels do; (iii) the probability of severe errors is lower in \lj models compared to the human expert re-score panels; (iv) the LLM Jury shows excellent agreement with primary expert panels' rankings. We find that the LLM jury combined with AI model diagnoses can be used to identify ward diagnoses at high risk of error, enabling targeted expert review and improved panel efficiency; (v) LLM jury models show no self-preference bias. They did not score diagnoses generated by their own underlying model or models from the same vendor more (or less) favourably than those generated by other models. Finally, we demonstrate that LLM jury calibration using isotonic regression improves alignment with human expert panel evaluations. Together, these results provide compelling evidence that a calibrated, multi-model LLM jury can serve as a trustworthy and reliable proxy for expert clinician evaluation in medical AI benchmarking.
format Preprint
id arxiv_https___arxiv_org_abs_2604_14892
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Can LLMs Score Medical Diagnoses and Clinical Reasoning as well as Expert Panels?
Rouillard, Amy
Mundia, Sitwala
Camara, Linda
Gramanie, Michael Cameron
Dangor, Ziyaad
Kalla, Ismail
Madhi, Shabir A.
Morar, Kajal
Ncube, Marlvin T.
Saloojee, Haroon
Bassett, Bruce A.
Machine Learning
Artificial Intelligence
Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators. Here, we evaluate an LLM jury composed of three frontier AI models scoring 3333 diagnoses on 300 real-world middle-income country (MIC) hospital cases. Model performance was benchmarked against expert clinician panel and independent human re-scoring panel evaluations. Both LLM and clinician-generated diagnoses are scored across four dimensions: diagnosis, differential diagnosis, clinical reasoning and negative treatment risk. For each of these, we assess scoring difference, inter-rater agreement, scoring stability, severe safety errors and the effect of post-hoc calibration. We find that: (i) the uncalibrated LLM jury scores are systematically lower than clinician panels scores; (ii) the LLM Jury preserves ordinal agreement and exhibits better concordance with the primary expert panels than the human expert re-score panels do; (iii) the probability of severe errors is lower in \lj models compared to the human expert re-score panels; (iv) the LLM Jury shows excellent agreement with primary expert panels' rankings. We find that the LLM jury combined with AI model diagnoses can be used to identify ward diagnoses at high risk of error, enabling targeted expert review and improved panel efficiency; (v) LLM jury models show no self-preference bias. They did not score diagnoses generated by their own underlying model or models from the same vendor more (or less) favourably than those generated by other models. Finally, we demonstrate that LLM jury calibration using isotonic regression improves alignment with human expert panel evaluations. Together, these results provide compelling evidence that a calibrated, multi-model LLM jury can serve as a trustworthy and reliable proxy for expert clinician evaluation in medical AI benchmarking.
title Can LLMs Score Medical Diagnoses and Clinical Reasoning as well as Expert Panels?
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2604.14892