Geometry-Calibrated Conformal Abstention for Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Rui, Chen, Yi, Xie, Sihong, Xiong, Hui
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914521132564480
author Xu, Rui
Chen, Yi
Xie, Sihong
Xiong, Hui
author_facet Xu, Rui
Chen, Yi
Xie, Sihong
Xiong, Hui
contents When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75 percent conditional correctness.
format Preprint
id arxiv_https___arxiv_org_abs_2604_27914
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Geometry-Calibrated Conformal Abstention for Language Models
Xu, Rui
Chen, Yi
Xie, Sihong
Xiong, Hui
Computation and Language
Machine Learning
When language models lack relevant knowledge for a given query, they frequently generate plausible responses that can be hallucinations, rather than admitting being agnostic about the answer. Retraining models to reward admitting ignorance can lead to overly conservative behaviors and poor generalization due to scarce evaluation benchmarks. We propose a post hoc framework, Conformal Abstention (CA), adapted from conformal prediction (CP) to determine whether to abstain from answering a query. CA provides finite-sample guarantees on both the probability of participation (i.e., not abstaining) and the probability that the generated response is correct. Importantly, the abstention decision relies on prediction confidence rather than the non-conformity scores used in CP, which are intractable for open-ended generation. To better align prediction confidence with the model's ignorance, we introduce a calibration strategy using representation geometry within the model to measure knowledge involvement in shaping the response. Experiments demonstrate that we improve selective answering significantly with 75 percent conditional correctness.
title Geometry-Calibrated Conformal Abstention for Language Models
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2604.27914