Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Xiaoyuan, Truong, Kimberly Le, Fogliato, Riccardo, Swamy, Gokul, Zhang, Weijian, Yang, Minglai, Ye, Longtian, Liu, Bangya, Liu, Minghao, Ilyas, Andrew, Wu, Steven
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918436284661760
author Zhu, Xiaoyuan
Truong, Kimberly Le
Fogliato, Riccardo
Swamy, Gokul
Zhang, Weijian
Yang, Minglai
Ye, Longtian
Liu, Bangya
Liu, Minghao
Ilyas, Andrew
Wu, Steven
author_facet Zhu, Xiaoyuan
Truong, Kimberly Le
Fogliato, Riccardo
Swamy, Gokul
Zhang, Weijian
Yang, Minglai
Ye, Longtian
Liu, Bangya
Liu, Minghao
Ilyas, Andrew
Wu, Steven
contents As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether these justifications help users distinguish correct answers from incorrect ones. We formalize this idea as error verifiability and propose $v_{\text{bal}}$, a balanced metric that measures whether justifications enable raters to accurately assess answer correctness, validated against human raters who show high agreement. We find that neither common approaches, such as post-training and model scaling, nor more targeted interventions recommended improve verifiability. We introduce two methods that succeed at improving verifiability: reflect-and-rephrase (RR) for mathematical reasoning and oracle-rephrase (OR) for factual QA, both of which improve verifiability by incorporating domain-appropriate external information. Together, our results establish error verifiability as a distinct dimension of response quality that does not emerge from accuracy improvements alone and requires dedicated, domain-aware methods to address.
format Preprint
id arxiv_https___arxiv_org_abs_2604_04418
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
Zhu, Xiaoyuan
Truong, Kimberly Le
Fogliato, Riccardo
Swamy, Gokul
Zhang, Weijian
Yang, Minglai
Ye, Longtian
Liu, Bangya
Liu, Minghao
Ilyas, Andrew
Wu, Steven
Human-Computer Interaction
Artificial Intelligence
As LLMs are deployed in high-stakes settings, users must judge the correctness of individual responses, often relying on model-generated justifications such as reasoning chains or explanations. Yet, no standard measure exists for whether these justifications help users distinguish correct answers from incorrect ones. We formalize this idea as error verifiability and propose $v_{\text{bal}}$, a balanced metric that measures whether justifications enable raters to accurately assess answer correctness, validated against human raters who show high agreement. We find that neither common approaches, such as post-training and model scaling, nor more targeted interventions recommended improve verifiability. We introduce two methods that succeed at improving verifiability: reflect-and-rephrase (RR) for mathematical reasoning and oracle-rephrase (OR) for factual QA, both of which improve verifiability by incorporating domain-appropriate external information. Together, our results establish error verifiability as a distinct dimension of response quality that does not emerge from accuracy improvements alone and requires dedicated, domain-aware methods to address.
title Justified or Just Convincing? Error Verifiability as a Dimension of LLM Quality
topic Human-Computer Interaction
Artificial Intelligence
url https://arxiv.org/abs/2604.04418