CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Kevin H., Yan, Chao, Baidya, Avinash, Brown, Katherine, Gao, Xiang, Xiong, Juming, Yin, Zhijun, Malin, Bradley A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917476445454336
author Guo, Kevin H.
Yan, Chao
Baidya, Avinash
Brown, Katherine
Gao, Xiang
Xiong, Juming
Yin, Zhijun
Malin, Bradley A.
author_facet Guo, Kevin H.
Yan, Chao
Baidya, Avinash
Brown, Katherine
Gao, Xiang
Xiong, Juming
Yin, Zhijun
Malin, Bradley A.
contents Medical large language model (LLM) evaluations rely on simplified, exam-style benchmarks that rarely reflect the ambiguity of real-world medical inquiries. We introduce the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework, which assesses how decision-space presentation, ambiguity, and uncertainty affect LLMs' reasoning on medical benchmarks. CLEAR systematically perturbs (1) the number of plausible answer options, (2) the presence of a ground truth or abstention option, and (3) the semantic framing of answer options. Applying CLEAR on three benchmarks evaluated across 17 LLMs reveals three notable limitations of existing evaluation methods. First, increasing the number of plausible answers degrades a model's ability to identify the correct answer and abstain against incorrect ones. Second, this lack of caution intensifies as the framing of abstention shifts from assertive rejection like "None of the Above" to uncertainty admission like "I don't know" (IDK). Notably, just including IDK in the answer space increases incorrect answer selections. Lastly, we formalize the performance gap between identifying the correct answer and abstaining from incorrect ones as the humility deficit, which worsens with model scale. Our findings reveal limitations in standard medical benchmarks and underscore that scaling alone does not resolve LLM reliability issues.
format Preprint
id arxiv_https___arxiv_org_abs_2605_01011
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
Guo, Kevin H.
Yan, Chao
Baidya, Avinash
Brown, Katherine
Gao, Xiang
Xiong, Juming
Yin, Zhijun
Malin, Bradley A.
Computation and Language
Artificial Intelligence
Machine Learning
Medical large language model (LLM) evaluations rely on simplified, exam-style benchmarks that rarely reflect the ambiguity of real-world medical inquiries. We introduce the CLinical Evaluation of Ambiguity and Reliability (CLEAR) framework, which assesses how decision-space presentation, ambiguity, and uncertainty affect LLMs' reasoning on medical benchmarks. CLEAR systematically perturbs (1) the number of plausible answer options, (2) the presence of a ground truth or abstention option, and (3) the semantic framing of answer options. Applying CLEAR on three benchmarks evaluated across 17 LLMs reveals three notable limitations of existing evaluation methods. First, increasing the number of plausible answers degrades a model's ability to identify the correct answer and abstain against incorrect ones. Second, this lack of caution intensifies as the framing of abstention shifts from assertive rejection like "None of the Above" to uncertainty admission like "I don't know" (IDK). Notably, just including IDK in the answer space increases incorrect answer selections. Lastly, we formalize the performance gap between identifying the correct answer and abstaining from incorrect ones as the humility deficit, which worsens with model scale. Our findings reveal limitations in standard medical benchmarks and underscore that scaling alone does not resolve LLM reliability issues.
title CLEAR: Revealing How Noise and Ambiguity Degrade Reliability in LLMs for Medicine
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2605.01011