Performance of Large Language Models in Supporting Medical Diagnosis and Treatment
Fuente:
arXiv
Saved in:
| Main Authors: | , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866915241877569536 |
|---|---|
| author | Sousa, Diogo Barbosa, Guilherme Rocha, Catarina Oliveira, Dulce |
| author_facet | Sousa, Diogo Barbosa, Guilherme Rocha, Catarina Oliveira, Dulce |
| contents | The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_10405 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Performance of Large Language Models in Supporting Medical Diagnosis and Treatment Sousa, Diogo Barbosa, Guilherme Rocha, Catarina Oliveira, Dulce Computation and Language Artificial Intelligence Emerging Technologies Human-Computer Interaction I.2.7; J.3 The integration of Large Language Models (LLMs) into healthcare holds significant potential to enhance diagnostic accuracy and support medical treatment planning. These AI-driven systems can analyze vast datasets, assisting clinicians in identifying diseases, recommending treatments, and predicting patient outcomes. This study evaluates the performance of a range of contemporary LLMs, including both open-source and closed-source models, on the 2024 Portuguese National Exam for medical specialty access (PNA), a standardized medical knowledge assessment. Our results highlight considerable variation in accuracy and cost-effectiveness, with several models demonstrating performance exceeding human benchmarks for medical students on this specific task. We identify leading models based on a combined score of accuracy and cost, discuss the implications of reasoning methodologies like Chain-of-Thought, and underscore the potential for LLMs to function as valuable complementary tools aiding medical professionals in complex clinical decision-making. |
| title | Performance of Large Language Models in Supporting Medical Diagnosis and Treatment |
| topic | Computation and Language Artificial Intelligence Emerging Technologies Human-Computer Interaction I.2.7; J.3 |
| url | https://arxiv.org/abs/2504.10405 |