MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Hao, Yuexing, Alhamoud, Kumail, Jeong, Hyewon, Zhang, Haoran, Puri, Isha, Torr, Philip, Schaekermann, Mike, Stern, Ariel D., Ghassemi, Marzyeh
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915313116774400
author Hao, Yuexing
Alhamoud, Kumail
Jeong, Hyewon
Zhang, Haoran
Puri, Isha
Torr, Philip
Schaekermann, Mike
Stern, Ariel D.
Ghassemi, Marzyeh
author_facet Hao, Yuexing
Alhamoud, Kumail
Jeong, Hyewon
Zhang, Haoran
Puri, Isha
Torr, Philip
Schaekermann, Mike
Stern, Ariel D.
Ghassemi, Marzyeh
contents Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may reach accurate conclusions through flawed processes. In this study, we introduce the MedPAIR (Medical Dataset Comparing Physicians and AI Relevance Estimation and Question Answering) dataset to evaluate how physician trainees and LLMs prioritize relevant information when answering QA questions. We obtain annotations on 1,300 QA pairs from 36 physician trainees, labeling each sentence within the question components for relevance. We compare these relevance estimates to those for LLMs, and further evaluate the impact of these "relevant" subsets on downstream task performance for both physician trainees and LLMs. We find that LLMs are frequently not aligned with the content relevance estimates of physician trainees. After filtering out physician trainee-labeled irrelevant sentences, accuracy improves for both the trainees and the LLMs. All LLM and physician trainee-labeled data are available at: http://medpair.csail.mit.edu/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_24040
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
Hao, Yuexing
Alhamoud, Kumail
Jeong, Hyewon
Zhang, Haoran
Puri, Isha
Torr, Philip
Schaekermann, Mike
Stern, Ariel D.
Ghassemi, Marzyeh
Computation and Language
Artificial Intelligence
Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may reach accurate conclusions through flawed processes. In this study, we introduce the MedPAIR (Medical Dataset Comparing Physicians and AI Relevance Estimation and Question Answering) dataset to evaluate how physician trainees and LLMs prioritize relevant information when answering QA questions. We obtain annotations on 1,300 QA pairs from 36 physician trainees, labeling each sentence within the question components for relevance. We compare these relevance estimates to those for LLMs, and further evaluate the impact of these "relevant" subsets on downstream task performance for both physician trainees and LLMs. We find that LLMs are frequently not aligned with the content relevance estimates of physician trainees. After filtering out physician trainee-labeled irrelevant sentences, accuracy improves for both the trainees and the LLMs. All LLM and physician trainee-labeled data are available at: http://medpair.csail.mit.edu/.
title MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.24040