Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Gourabathina, Abinitha, Padhi, Inkit, Nagireddy, Manish, Chaudhury, Subhajit, Sattigeri, Prasanna
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914440587247616
author Gourabathina, Abinitha
Padhi, Inkit
Nagireddy, Manish
Chaudhury, Subhajit
Sattigeri, Prasanna
author_facet Gourabathina, Abinitha
Padhi, Inkit
Nagireddy, Manish
Chaudhury, Subhajit
Sattigeri, Prasanna
contents For Large Language Models (LLMs) to be reliably deployed, models must effectively know when not to answer: abstain. Reasoning models, in particular, have gained attention for impressive performance on complex tasks. However, reasoning models have been shown to have worse abstention abilities. Taking the vulnerabilities of reasoning models into account, we propose our Query Misalignment Framework. Hallucinations resulting in failed abstention can be reinterpreted as LLMs answering the wrong question (rather than answering a question incorrectly). Based on this framework, we develop a new class of state-of-the-art abstention methods called Trace Inversion. First, we generate the reasoning trace of a model. Based on only the trace, we then reconstruct the most likely query that the model responded to. Finally, we compare the initial query with the reconstructed query. Low similarity score between the initial query and reconstructed query suggests that the model likely answered the question incorrectly and is flagged to abstain. Extensive experiments demonstrate that Trace Inversion effectively boosts abstention performance in four frontier LLMs across nine abstention QA datasets, beating competitive baselines in 33 out of 36 settings.
format Preprint
id arxiv_https___arxiv_org_abs_2604_02230
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
Gourabathina, Abinitha
Padhi, Inkit
Nagireddy, Manish
Chaudhury, Subhajit
Sattigeri, Prasanna
Artificial Intelligence
For Large Language Models (LLMs) to be reliably deployed, models must effectively know when not to answer: abstain. Reasoning models, in particular, have gained attention for impressive performance on complex tasks. However, reasoning models have been shown to have worse abstention abilities. Taking the vulnerabilities of reasoning models into account, we propose our Query Misalignment Framework. Hallucinations resulting in failed abstention can be reinterpreted as LLMs answering the wrong question (rather than answering a question incorrectly). Based on this framework, we develop a new class of state-of-the-art abstention methods called Trace Inversion. First, we generate the reasoning trace of a model. Based on only the trace, we then reconstruct the most likely query that the model responded to. Finally, we compare the initial query with the reconstructed query. Low similarity score between the initial query and reconstructed query suggests that the model likely answered the question incorrectly and is flagged to abstain. Extensive experiments demonstrate that Trace Inversion effectively boosts abstention performance in four frontier LLMs across nine abstention QA datasets, beating competitive baselines in 33 out of 36 settings.
title Answering the Wrong Question: Reasoning Trace Inversion for Abstention in LLMs
topic Artificial Intelligence
url https://arxiv.org/abs/2604.02230