Enregistré dans:
Détails bibliographiques
Auteurs principaux: Yang, Haoyan, Shirkavand, Reza, Jin, Yukai, Zhou, Jiawei, Gao, Shangqian, Huang, Heng
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:https://arxiv.org/abs/2606.00251
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918532519821312
author Yang, Haoyan
Shirkavand, Reza
Jin, Yukai
Zhou, Jiawei
Gao, Shangqian
Huang, Heng
author_facet Yang, Haoyan
Shirkavand, Reza
Jin, Yukai
Zhou, Jiawei
Gao, Shangqian
Huang, Heng
contents The ability to recognize one's own limitations and decide whether to solve a problem or delegate is fundamental for reliable intelligent systems. Yet we show that modern large language models systematically lack this ability: across diverse model families and scales, they overestimate their competence and attempt queries they cannot solve. We refer to this ability as Capability Self-Assessment (CSA) and formulate it as a policy-learning problem, aiming to improve self-assessment while preserving the model's original capabilities. Our results show that reinforcement learning teaches CSA effectively, significantly outperforming supervised fine-tuning while preserving original capabilities. In contrast, supervised fine-tuning severely degrades the capabilities the model is meant to assess. Moreover, learned self-assessment behavior generalizes well out of distribution, suggesting that CSA is a transferable model trait. Finally, CSA is practically useful: it improves local-cloud decision making at inference time and provides a signal for targeted data selection during training.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00251
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Capability Self-Assessment: Teaching LLMs to Know Their Limits
Yang, Haoyan
Shirkavand, Reza
Jin, Yukai
Zhou, Jiawei
Gao, Shangqian
Huang, Heng
Artificial Intelligence
The ability to recognize one's own limitations and decide whether to solve a problem or delegate is fundamental for reliable intelligent systems. Yet we show that modern large language models systematically lack this ability: across diverse model families and scales, they overestimate their competence and attempt queries they cannot solve. We refer to this ability as Capability Self-Assessment (CSA) and formulate it as a policy-learning problem, aiming to improve self-assessment while preserving the model's original capabilities. Our results show that reinforcement learning teaches CSA effectively, significantly outperforming supervised fine-tuning while preserving original capabilities. In contrast, supervised fine-tuning severely degrades the capabilities the model is meant to assess. Moreover, learned self-assessment behavior generalizes well out of distribution, suggesting that CSA is a transferable model trait. Finally, CSA is practically useful: it improves local-cloud decision making at inference time and provides a signal for targeted data selection during training.
title Capability Self-Assessment: Teaching LLMs to Know Their Limits
topic Artificial Intelligence
url https://arxiv.org/abs/2606.00251