Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Peng, Dengyun, Chen, Qiguang, Liu, Bofei, Guan, Jiannan, Qin, Libo, Yan, Zheng, Liu, Jinhao, Zhang, Jianshu, Che, Wanxiang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914297603424256
author Peng, Dengyun
Chen, Qiguang
Liu, Bofei
Guan, Jiannan
Qin, Libo
Yan, Zheng
Liu, Jinhao
Zhang, Jianshu
Che, Wanxiang
author_facet Peng, Dengyun
Chen, Qiguang
Liu, Bofei
Guan, Jiannan
Qin, Libo
Yan, Zheng
Liu, Jinhao
Zhang, Jianshu
Che, Wanxiang
contents Ensuring large language model (LLM) reliability requires distinguishing objective unsolvability (inherent contradictions) from subjective capability limitations (tasks exceeding model competence). Current LLMs often conflate these dimensions, leading to hallucinations in which they return confident answers to inherently unsolvable queries. To address this issue, we propose a multi-domain dataset containing both solvable and unsolvable questions, UnsolvableQA, together with an alignment framework, UnsolvableRL. First, we construct UnsolvableQA by "Reverse Construction" that systematically injects logical contradictions into otherwise valid reasoning chains. Second, we introduce UnsolvableRL, a reinforcement learning paradigm that balances objective unsolvability detection with calibrated confidence under capability limits. Empirically, our approach achieves robust unsolvability detection (>85% detection rate) and boosts solvable reasoning accuracy from 43.4% to 69.4% on Qwen3-4B-Instruct. Crucially, we identify a data-training interaction: strict alignment constraints induce Capability Collapse without unsolvable data, but act as a regularizer for rigor when such data are included, thereby improving overall robustness. Our code and data are available at https://github.com/sfasfaffa/unsolvableQA .
format Preprint
id arxiv_https___arxiv_org_abs_2512_01661
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
Peng, Dengyun
Chen, Qiguang
Liu, Bofei
Guan, Jiannan
Qin, Libo
Yan, Zheng
Liu, Jinhao
Zhang, Jianshu
Che, Wanxiang
Computation and Language
Artificial Intelligence
Ensuring large language model (LLM) reliability requires distinguishing objective unsolvability (inherent contradictions) from subjective capability limitations (tasks exceeding model competence). Current LLMs often conflate these dimensions, leading to hallucinations in which they return confident answers to inherently unsolvable queries. To address this issue, we propose a multi-domain dataset containing both solvable and unsolvable questions, UnsolvableQA, together with an alignment framework, UnsolvableRL. First, we construct UnsolvableQA by "Reverse Construction" that systematically injects logical contradictions into otherwise valid reasoning chains. Second, we introduce UnsolvableRL, a reinforcement learning paradigm that balances objective unsolvability detection with calibrated confidence under capability limits. Empirically, our approach achieves robust unsolvability detection (>85% detection rate) and boosts solvable reasoning accuracy from 43.4% to 69.4% on Qwen3-4B-Instruct. Crucially, we identify a data-training interaction: strict alignment constraints induce Capability Collapse without unsolvable data, but act as a regularizer for rigor when such data are included, thereby improving overall robustness. Our code and data are available at https://github.com/sfasfaffa/unsolvableQA .
title Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2512.01661