Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty
Fuente:
arXiv
Saved in:
| Main Authors: | Machcha, Sravanthi, Yerra, Sushrita, Gupta, Sahil, Sahoo, Aishwarya, Sultana, Sharmin, Yu, Hong, Yao, Zonghai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025)
by: Tran, Hieu, et al.
Published: (2025)
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026)
by: Zhou, Xinyu, et al.
Published: (2026)
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
by: Rakshit, Sushrita, et al.
Published: (2026)
by: Rakshit, Sushrita, et al.
Published: (2026)
EHR Interaction Between Patients and AI: NoteAid EHR Interaction
by: Zhang, Xiaocheng, et al.
Published: (2023)
by: Zhang, Xiaocheng, et al.
Published: (2023)
KnowRL: Teaching Language Models to Know What They Know
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
Reasoning about Uncertainty: Do Reasoning Models Know When They Don't Know?
by: Mei, Zhiting, et al.
Published: (2025)
by: Mei, Zhiting, et al.
Published: (2025)
MedCOD: Enhancing English-to-Spanish Medical Translation of Large Language Models Using Enriched Chain-of-Dictionary Framework
by: Salim, Md Shahidul, et al.
Published: (2025)
by: Salim, Md Shahidul, et al.
Published: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
by: Sahoo, Subramanyam
Published: (2026)
by: Sahoo, Subramanyam
Published: (2026)
BioInstruct: Instruction Tuning of Large Language Models for Biomedical Natural Language Processing
by: Tran, Hieu, et al.
Published: (2023)
by: Tran, Hieu, et al.
Published: (2023)
ReadCtrl: Personalizing text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2024)
by: Tran, Hieu, et al.
Published: (2024)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
by: Sun, Yuxi, et al.
Published: (2025)
by: Sun, Yuxi, et al.
Published: (2025)
Teaching LLMs to Abstain via Fine-Grained Semantic Confidence Reward
by: An, Hao, et al.
Published: (2025)
by: An, Hao, et al.
Published: (2025)
Know When to Abstain: Optimal Selective Classification with Likelihood Ratios
by: Heng, Alvin, et al.
Published: (2025)
by: Heng, Alvin, et al.
Published: (2025)
Enhancing LLMs for Identifying and Prioritizing Important Medical Jargons from Electronic Health Record Notes Utilizing Data Augmentation
by: Jang, Won Seok, et al.
Published: (2025)
by: Jang, Won Seok, et al.
Published: (2025)
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization
by: Mishra, Prakamya, et al.
Published: (2024)
by: Mishra, Prakamya, et al.
Published: (2024)
From Scores to Steps: Diagnosing and Improving LLM Performance in Evidence-Based Medical Calculations
by: Wang, Benlu, et al.
Published: (2025)
by: Wang, Benlu, et al.
Published: (2025)
README: Bridging Medical Jargon and Lay Understanding for Patient Education through Data-Centric NLP
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
Uncertainty Estimation of Large Language Models in Medical Question Answering
by: Wu, Jiaxin, et al.
Published: (2024)
by: Wu, Jiaxin, et al.
Published: (2024)
Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs
by: Kale, Sahil
Published: (2025)
by: Kale, Sahil
Published: (2025)
Mind the Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering
by: Liu, Yaokun, et al.
Published: (2026)
by: Liu, Yaokun, et al.
Published: (2026)
MedQA-CS: Objective Structured Clinical Examination (OSCE)-Style Benchmark for Evaluating LLM Clinical Skills
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
Do Retrieval Augmented Language Models Know When They Don't Know?
by: Zhou, Youchao, et al.
Published: (2025)
by: Zhou, Youchao, et al.
Published: (2025)
DischargeSim: A Simulation Benchmark for Educational Doctor-Patient Communication at Discharge
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
Do not Abstain! Identify and Solve the Uncertainty
by: Liu, Jingyu, et al.
Published: (2025)
by: Liu, Jingyu, et al.
Published: (2025)
Text Knows What, Tables Know When: Clinical Timeline Reconstruction via Retrieval-Augmented Multimodal Alignment
by: Kumar, Sayantan, et al.
Published: (2026)
by: Kumar, Sayantan, et al.
Published: (2026)
When Models Know More Than They Say: Probing Analogical Reasoning in LLMs
by: McGovern, Hope, et al.
Published: (2026)
by: McGovern, Hope, et al.
Published: (2026)
Surgical Feature-Space Decomposition of LLMs: Why, When and How?
by: Chavan, Arnav, et al.
Published: (2024)
by: Chavan, Arnav, et al.
Published: (2024)
Mapping Clinical Doubt: Locating Linguistic Uncertainty in LLMs
by: Sridhar, Srivarshinee, et al.
Published: (2025)
by: Sridhar, Srivarshinee, et al.
Published: (2025)
What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"
by: Lee, Joosung, et al.
Published: (2026)
by: Lee, Joosung, et al.
Published: (2026)
MCQG-SRefine: Multiple Choice Question Generation and Evaluation with Iterative Self-Critique, Correction, and Comparison Feedback
by: Yao, Zonghai, et al.
Published: (2024)
by: Yao, Zonghai, et al.
Published: (2024)
SynthEHR-Eviction: Enhancing Eviction SDoH Detection with LLM-Augmented Synthetic EHR Data
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
by: Zhai, Skylar, et al.
Published: (2026)
by: Zhai, Skylar, et al.
Published: (2026)
Emotionally-Aware Agents for Dispute Resolution
by: Rakshit, Sushrita, et al.
Published: (2025)
by: Rakshit, Sushrita, et al.
Published: (2025)
Question Answering on Patient Medical Records with Private Fine-Tuned LLMs
by: Kothari, Sara, et al.
Published: (2025)
by: Kothari, Sara, et al.
Published: (2025)
Clarify, Abstain or Answer? Strategising in Conversation with Belief-Augmented Generation
by: Baan, Joris, et al.
Published: (2026)
by: Baan, Joris, et al.
Published: (2026)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
CaRT: Teaching LLM Agents to Know When They Know Enough
by: Liu, Grace, et al.
Published: (2025)
by: Liu, Grace, et al.
Published: (2025)
ChatCLIDS: Simulating Persuasive AI Dialogues to Promote Closed-Loop Insulin Adoption in Type 1 Diabetes Care
by: Yao, Zonghai, et al.
Published: (2025)
by: Yao, Zonghai, et al.
Published: (2025)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
by: Kale, Sahil, et al.
Published: (2025)
by: Kale, Sahil, et al.
Published: (2025)
Similar Items
-
Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
by: Yao, Zonghai, et al.
Published: (2023) -
MedReadCtrl: Personalizing medical text generation with readability-controlled instruction learning
by: Tran, Hieu, et al.
Published: (2025) -
When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond?
by: Zhou, Xinyu, et al.
Published: (2026) -
Pseudo-Deliberation in Language Models: When Reasoning Fails to Align Values and Actions
by: Rakshit, Sushrita, et al.
Published: (2026) -
EHR Interaction Between Patients and AI: NoteAid EHR Interaction
by: Zhang, Xiaocheng, et al.
Published: (2023)