Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Machcha, Sravanthi, Yerra, Sushrita, Gupta, Sahil, Sahoo, Aishwarya, Sultana, Sharmin, Yu, Hong, Yao, Zonghai
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909997481328640
author Machcha, Sravanthi
Yerra, Sushrita
Gupta, Sahil
Sahoo, Aishwarya
Sultana, Sharmin
Yu, Hong
Yao, Zonghai
author_facet Machcha, Sravanthi
Yerra, Sushrita
Gupta, Sahil
Sahoo, Aishwarya
Sultana, Sharmin
Yu, Hong
Yao, Zonghai
contents Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.
format Preprint
id arxiv_https___arxiv_org_abs_2601_12471
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
Machcha, Sravanthi
Yerra, Sushrita
Gupta, Sahil
Sahoo, Aishwarya
Sultana, Sharmin
Yu, Hong
Yao, Zonghai
Computation and Language
Artificial Intelligence
Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.
title Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2601.12471