Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Hanjie, Fang, Zhouxiang, Singla, Yash, Dredze, Mark
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915804215246848
author Chen, Hanjie
Fang, Zhouxiang
Singla, Yash
Dredze, Mark
author_facet Chen, Hanjie
Fang, Zhouxiang
Singla, Yash
Dredze, Mark
contents LLMs have demonstrated impressive performance in answering medical questions, such as achieving passing scores on medical licensing examinations. However, medical board exams or general clinical questions do not capture the complexity of realistic clinical cases. Moreover, the lack of reference explanations means we cannot easily evaluate the reasoning of model decisions, a crucial component of supporting doctors in making complex medical decisions. To address these challenges, we construct two new datasets: JAMA Clinical Challenge and Medbullets. Datasets and code are available at https://github.com/HanjieChen/ChallengeClinicalQA. JAMA Clinical Challenge consists of questions based on challenging clinical cases, while Medbullets comprises simulated clinical questions. Both datasets are structured as multiple-choice question-answering tasks, accompanied by expert-written explanations. We evaluate seven LLMs on the two datasets using various prompts. Experiments demonstrate that our datasets are harder than previous benchmarks. In-depth automatic and human evaluations of model-generated explanations provide insights into the promise and deficiency of LLMs for explainable medical QA.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18060
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
Chen, Hanjie
Fang, Zhouxiang
Singla, Yash
Dredze, Mark
Computation and Language
LLMs have demonstrated impressive performance in answering medical questions, such as achieving passing scores on medical licensing examinations. However, medical board exams or general clinical questions do not capture the complexity of realistic clinical cases. Moreover, the lack of reference explanations means we cannot easily evaluate the reasoning of model decisions, a crucial component of supporting doctors in making complex medical decisions. To address these challenges, we construct two new datasets: JAMA Clinical Challenge and Medbullets. Datasets and code are available at https://github.com/HanjieChen/ChallengeClinicalQA. JAMA Clinical Challenge consists of questions based on challenging clinical cases, while Medbullets comprises simulated clinical questions. Both datasets are structured as multiple-choice question-answering tasks, accompanied by expert-written explanations. We evaluate seven LLMs on the two datasets using various prompts. Experiments demonstrate that our datasets are harder than previous benchmarks. In-depth automatic and human evaluations of model-generated explanations provide insights into the promise and deficiency of LLMs for explainable medical QA.
title Benchmarking Large Language Models on Answering and Explaining Challenging Medical Questions
topic Computation and Language
url https://arxiv.org/abs/2402.18060