Performance of Large Language Models in Supporting Medical Diagnosis and Treatment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sousa, Diogo, Barbosa, Guilherme, Rocha, Catarina, Oliveira, Dulce
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915241877569536
author Sousa, Diogo
Barbosa, Guilherme
Rocha, Catarina
Oliveira, Dulce
author_facet Sousa, Diogo
Barbosa, Guilherme
Rocha, Catarina
Oliveira, Dulce
contents The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10405
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Performance of Large Language Models in Supporting Medical Diagnosis and Treatment
Sousa, Diogo
Barbosa, Guilherme
Rocha, Catarina
Oliveira, Dulce
Computation and Language
Artificial Intelligence
Emerging Technologies
Human-Computer Interaction
I.2.7; J.3
The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making.
title Performance of Large Language Models in Supporting Medical Diagnosis and Treatment
topic Computation and Language
Artificial Intelligence
Emerging Technologies
Human-Computer Interaction
I.2.7; J.3
url https://arxiv.org/abs/2504.10405