AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
Fuente:
arXiv
Saved in:
| Main Authors: | Kirichenko, Polina, Ibrahim, Mark, Chaudhuri, Kamalika, Bell, Samuel J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
by: Tomani, Christian, et al.
Published: (2024)
by: Tomani, Christian, et al.
Published: (2024)
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026)
by: Gourabathina, Abinitha, et al.
Published: (2026)
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026)
by: Wang, Erchi, et al.
Published: (2026)
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)
by: Akter, Syeda Nahida, et al.
Published: (2024)
Benchmarking Visual LLMs Resilience to Unanswerable Questions on Visually Rich Documents
by: Napolitano, Davide, et al.
Published: (2025)
by: Napolitano, Davide, et al.
Published: (2025)
I Could've Asked That: Reformulating Unanswerable Questions
by: Zhao, Wenting, et al.
Published: (2024)
by: Zhao, Wenting, et al.
Published: (2024)
Towards Unbiased Evaluation of Detecting Unanswerable Questions in EHRSQL
by: Yang, Yongjin, et al.
Published: (2024)
by: Yang, Yongjin, et al.
Published: (2024)
MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
by: Zhu, Yanxu, et al.
Published: (2025)
by: Zhu, Yanxu, et al.
Published: (2025)
ClinDet-Bench: Beyond Abstention, Evaluating Judgment Determinability of LLMs in Clinical Decision-Making
by: Watanabe, Yusuke, et al.
Published: (2026)
by: Watanabe, Yusuke, et al.
Published: (2026)
Influence-based Attributions can be Manipulated
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
CausalAbstain: Enhancing Multilingual LLMs with Causal Reasoning for Trustworthy Abstention
by: Sun, Yuxi, et al.
Published: (2025)
by: Sun, Yuxi, et al.
Published: (2025)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
Knowing When Not to Answer: Abstention-Aware Scientific Reasoning
by: Abdaljalil, Samir, et al.
Published: (2026)
by: Abdaljalil, Samir, et al.
Published: (2026)
Explicit Abstention Knobs for Predictable Reliability in Video Question Answering
by: Ortiz, Jorge
Published: (2025)
by: Ortiz, Jorge
Published: (2025)
Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions
by: Atif, Farah, et al.
Published: (2025)
by: Atif, Farah, et al.
Published: (2025)
InductionBench: LLMs Fail in the Simplest Complexity Class
by: Hua, Wenyue, et al.
Published: (2025)
by: Hua, Wenyue, et al.
Published: (2025)
Privacy-Preserving Retrieval-Augmented Generation with Differential Privacy
by: Koga, Tatsuki, et al.
Published: (2024)
by: Koga, Tatsuki, et al.
Published: (2024)
LLM Knowledge is Brittle: Truthfulness Representations Rely on Superficial Resemblance
by: Haller, Patrick, et al.
Published: (2025)
by: Haller, Patrick, et al.
Published: (2025)
WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
by: Evtimov, Ivan, et al.
Published: (2025)
by: Evtimov, Ivan, et al.
Published: (2025)
ReEfBench: Quantifying the Reasoning Efficiency of LLMs
by: Fu, Zhizhang, et al.
Published: (2026)
by: Fu, Zhizhang, et al.
Published: (2026)
Cost-Saving LLM Cascades with Early Abstention
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
AgentDAM: Privacy Leakage Evaluation for Autonomous Web Agents
by: Zharmagambetov, Arman, et al.
Published: (2025)
by: Zharmagambetov, Arman, et al.
Published: (2025)
ExpProof : Operationalizing Explanations for Confidential Models with ZKPs
by: Yadav, Chhavi, et al.
Published: (2025)
by: Yadav, Chhavi, et al.
Published: (2025)
FairProof : Confidential and Certifiable Fairness for Neural Networks
by: Yadav, Chhavi, et al.
Published: (2024)
by: Yadav, Chhavi, et al.
Published: (2024)
FeynmanBench: Benchmarking Multimodal LLMs on Diagrammatic Physics Reasoning
by: Wang, Zeyu, et al.
Published: (2026)
by: Wang, Zeyu, et al.
Published: (2026)
FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
by: Zhang, Qian-Wen, et al.
Published: (2025)
by: Zhang, Qian-Wen, et al.
Published: (2025)
Fail Fast, or Ask: Mitigating the Deficiencies of Reasoning LLMs with Human-in-the-Loop Systems Engineering
by: Zellinger, Michael J., et al.
Published: (2025)
by: Zellinger, Michael J., et al.
Published: (2025)
ReasonCACHE: Teaching LLMs To Reason Without Weight Updates
by: Gupta, Sharut, et al.
Published: (2026)
by: Gupta, Sharut, et al.
Published: (2026)
GeoGramBench: Benchmarking the Geometric Program Reasoning in Modern LLMs
by: Luo, Shixian, et al.
Published: (2025)
by: Luo, Shixian, et al.
Published: (2025)
Reasoning Fails Where Step Flow Breaks
by: Xu, Xiaoyu, et al.
Published: (2026)
by: Xu, Xiaoyu, et al.
Published: (2026)
TopoBench: Benchmarking LLMs on Hard Topological Reasoning
by: Maniparambil, Mayug, et al.
Published: (2026)
by: Maniparambil, Mayug, et al.
Published: (2026)
Reliable Text-to-SQL with Adaptive Abstention
by: Chen, Kaiwen, et al.
Published: (2025)
by: Chen, Kaiwen, et al.
Published: (2025)
Safety Alignment of LMs via Non-cooperative Games
by: Paulus, Anselm, et al.
Published: (2025)
by: Paulus, Anselm, et al.
Published: (2025)
LatentRefusal: Latent-Signal Refusal for Unanswerable Text-to-SQL Queries
by: Ren, Xuancheng, et al.
Published: (2026)
by: Ren, Xuancheng, et al.
Published: (2026)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?
by: Yao, Jianzhu, et al.
Published: (2025)
by: Yao, Jianzhu, et al.
Published: (2025)
Tracking the Limits of Knowledge Propagation: How LLMs Fail at Multi-Step Reasoning with Conflicting Knowledge
by: Feng, Yiyang, et al.
Published: (2026)
by: Feng, Yiyang, et al.
Published: (2026)
DW-Bench: Benchmarking LLMs on Data Warehouse Graph Topology Reasoning
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
by: Ahmed, Ahmed G. A. H, et al.
Published: (2026)
HiBench: Benchmarking LLMs Capability on Hierarchical Structure Reasoning
by: Jiang, Zhuohang, et al.
Published: (2025)
by: Jiang, Zhuohang, et al.
Published: (2025)
Similar Items
-
Uncertainty-Based Abstention in LLMs Improves Safety and Reduces Hallucinations
by: Tomani, Christian, et al.
Published: (2024) -
Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
by: Gourabathina, Abinitha, et al.
Published: (2026) -
Answering the Unanswerable Is to Err Knowingly: Analyzing and Mitigating Abstention Failures in Large Reasoning Models
by: Liu, Yi, et al.
Published: (2025) -
DPrivBench: Benchmarking LLMs' Reasoning for Differential Privacy
by: Wang, Erchi, et al.
Published: (2026) -
VISREAS: Complex Visual Reasoning with Unanswerable Questions
by: Akter, Syeda Nahida, et al.
Published: (2024)