Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Tianyi, Xu, Samuel, Dang, Jason Tansong, Yan, Samuel, Yin, Kimberley
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918507561615360
author Huang, Tianyi
Xu, Samuel
Dang, Jason Tansong
Yan, Samuel
Yin, Kimberley
author_facet Huang, Tianyi
Xu, Samuel
Dang, Jason Tansong
Yan, Samuel
Yin, Kimberley
contents Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what the evidence supports. We study this failure mode as overcommitment control and introduce compositional selective specificity (CSS), a post-generation layer that decomposes an answer into claims, proposes coarser backoffs, and emits each claim at the most specific calibrated level that appears admissible. The method is designed to express uncertainty as a local semantic backoff rather than as a whole-answer refusal. Across a full LongFact run and HotpotQA pilots, calibrated CSS improves the risk-utility trade-off of fixed drafts. On the full LongFact run, it raises overcommitment-aware utility from 0.846 to 0.913 relative to the no-CSS output while achieving 0.938 specificity retention. These results suggest that claim-level specificity control is a useful uncertainty interface for agentic systems and a target for future distribution-free validity layers.
format Preprint
id arxiv_https___arxiv_org_abs_2604_17487
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
Huang, Tianyi
Xu, Samuel
Dang, Jason Tansong
Yan, Samuel
Yin, Kimberley
Computation and Language
Agentic systems often fail not by being entirely wrong, but by being too precise: a response may be generally useful while particular claims exceed what the evidence supports. We study this failure mode as overcommitment control and introduce compositional selective specificity (CSS), a post-generation layer that decomposes an answer into claims, proposes coarser backoffs, and emits each claim at the most specific calibrated level that appears admissible. The method is designed to express uncertainty as a local semantic backoff rather than as a whole-answer refusal. Across a full LongFact run and HotpotQA pilots, calibrated CSS improves the risk-utility trade-off of fixed drafts. On the full LongFact run, it raises overcommitment-aware utility from 0.846 to 0.913 relative to the no-CSS output while achieving 0.938 specificity retention. These results suggest that claim-level specificity control is a useful uncertainty interface for agentic systems and a target for future distribution-free validity layers.
title Answer Only as Precisely as Justified: Calibrated Claim-Level Specificity Control for Agentic Systems
topic Computation and Language
url https://arxiv.org/abs/2604.17487