Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhai, Skylar, Liang, Jingcheng, Kang, Dongyeop
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918453272641536
author Zhai, Skylar
Liang, Jingcheng
Kang, Dongyeop
author_facet Zhai, Skylar
Liang, Jingcheng
Kang, Dongyeop
contents Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missing information. Existing abstention methods either train models to produce generic refusals or encourage follow-up clarifications without verifying whether those clarifications identify the key missing information. We study queries that are clear in meaning but cannot be reliably resolved from the given information, and argue that a reliable model should not only abstain, but also explain what is missing. We propose a clarification-aware RLVR reward that, while rewarding correct answers on answerable queries, jointly optimizes explicit abstention and semantically aligned post-refusal clarification on unanswerable queries. Using this reward, we train Abstain-R1, a 3B model that improves abstention and clarification on unanswerable queries while preserving strong performance on answerable ones. Experiments on Abstain-Test, Abstain-QA, and SelfAware show that Abstain-R1 substantially improves over its base model and achieves unanswerable-query behavior competitive with larger systems including DeepSeek-R1, suggesting that calibrated abstention and clarification can be learned through verifiable rewards rather than emerging from scale alone.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17073
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
Zhai, Skylar
Liang, Jingcheng
Kang, Dongyeop
Computation and Language
Artificial Intelligence
Reinforcement fine-tuning improves the reasoning ability of large language models, but it can also encourage them to answer unanswerable queries by guessing or hallucinating missing information. Existing abstention methods either train models to produce generic refusals or encourage follow-up clarifications without verifying whether those clarifications identify the key missing information. We study queries that are clear in meaning but cannot be reliably resolved from the given information, and argue that a reliable model should not only abstain, but also explain what is missing. We propose a clarification-aware RLVR reward that, while rewarding correct answers on answerable queries, jointly optimizes explicit abstention and semantically aligned post-refusal clarification on unanswerable queries. Using this reward, we train Abstain-R1, a 3B model that improves abstention and clarification on unanswerable queries while preserving strong performance on answerable ones. Experiments on Abstain-Test, Abstain-QA, and SelfAware show that Abstain-R1 substantially improves over its base model and achieves unanswerable-query behavior competitive with larger systems including DeepSeek-R1, suggesting that calibrated abstention and clarification can be learned through verifiable rewards rather than emerging from scale alone.
title Abstain-R1: Calibrated Abstention and Post-Refusal Clarification via Verifiable RL
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2604.17073