Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Adib, Shefayat E Shams, Sani, Ahmed Alfey, Esham, Ekramul Alam, Abrar, Ajwad, Chowdhury, Tareque Mohmud
Formato: Preprint
Publicado: 2026
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910023531102208
author Adib, Shefayat E Shams
Sani, Ahmed Alfey
Esham, Ekramul Alam
Abrar, Ajwad
Chowdhury, Tareque Mohmud
author_facet Adib, Shefayat E Shams
Sani, Ahmed Alfey
Esham, Ekramul Alam
Abrar, Ajwad
Chowdhury, Tareque Mohmud
contents Recently, Large Language Models (LLMs) have gained significant traction in medical domain, especially in developing a QA systems to Medical QA systems for enhancing access to healthcare in low-resourced settings. This paper compares five LLMs deployed between April 2024 and August 2025 for medical QA, using the iCliniq dataset, containing 38,000 medical questions and answers of diverse specialties. Our models include Llama-3-8B-Instruct, Llama 3.2 3B, Llama 3.3 70B Instruct, Llama-4-Maverick-17B-128E-Instruct, and GPT-5-mini. We are using a zero-shot evaluation methodology and using BLEU and ROUGE metrics to evaluate performance without specialized fine-tuning. Our results show that larger models like Llama 3.3 70B Instruct outperform smaller models, consistent with observed scaling benefits in clinical tasks. It is notable that, Llama-4-Maverick-17B exhibited more competitive results, thus highlighting evasion efficiency trade-offs relevant for practical deployment. These findings align with advancements in LLM capabilities toward professional-level medical reasoning and reflect the increasing feasibility of LLM-supported QA systems in the real clinical environments. This benchmark aims to serve as a standardized setting for future study to minimize model size, computational resources and to maximize clinical utility in medical NLP applications.
format Preprint
id arxiv_https___arxiv_org_abs_2602_14564
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
Adib, Shefayat E Shams
Sani, Ahmed Alfey
Esham, Ekramul Alam
Abrar, Ajwad
Chowdhury, Tareque Mohmud
Computation and Language
Recently, Large Language Models (LLMs) have gained significant traction in medical domain, especially in developing a QA systems to Medical QA systems for enhancing access to healthcare in low-resourced settings. This paper compares five LLMs deployed between April 2024 and August 2025 for medical QA, using the iCliniq dataset, containing 38,000 medical questions and answers of diverse specialties. Our models include Llama-3-8B-Instruct, Llama 3.2 3B, Llama 3.3 70B Instruct, Llama-4-Maverick-17B-128E-Instruct, and GPT-5-mini. We are using a zero-shot evaluation methodology and using BLEU and ROUGE metrics to evaluate performance without specialized fine-tuning. Our results show that larger models like Llama 3.3 70B Instruct outperform smaller models, consistent with observed scaling benefits in clinical tasks. It is notable that, Llama-4-Maverick-17B exhibited more competitive results, thus highlighting evasion efficiency trade-offs relevant for practical deployment. These findings align with advancements in LLM capabilities toward professional-level medical reasoning and reflect the increasing feasibility of LLM-supported QA systems in the real clinical environments. This benchmark aims to serve as a standardized setting for future study to minimize model size, computational resources and to maximize clinical utility in medical NLP applications.
title Assessing Large Language Models for Medical QA: Zero-Shot and LLM-as-a-Judge Evaluation
topic Computation and Language
url https://arxiv.org/abs/2602.14564