A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Lingjun, Daumé III, Hal
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908562489344000
author Zhao, Lingjun
Daumé III, Hal
author_facet Zhao, Lingjun
Daumé III, Hal
contents Faithful free-text explanations are important to ensure transparency in high-stakes AI decision-making contexts, but they are challenging to generate by language models and assess by humans. In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of weight of evidence. This measure quantifies how much a free-text explanation supports or opposes a prediction, serving as an important aspect of explanation faithfulness. Our analysis reveals that more than 62% explanations generated by large language models lack this consistency. We show that applying direct preference optimization improves the consistency of generated explanations across three model families, with improvement ranging from 43.1% to 292.3%. Furthermore, we demonstrate that optimizing this consistency measure can improve explanation faithfulness by up to 9.7%.
format Preprint
id arxiv_https___arxiv_org_abs_2505_19299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
Zhao, Lingjun
Daumé III, Hal
Computation and Language
Artificial Intelligence
Faithful free-text explanations are important to ensure transparency in high-stakes AI decision-making contexts, but they are challenging to generate by language models and assess by humans. In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of weight of evidence. This measure quantifies how much a free-text explanation supports or opposes a prediction, serving as an important aspect of explanation faithfulness. Our analysis reveals that more than 62% explanations generated by large language models lack this consistency. We show that applying direct preference optimization improves the consistency of generated explanations across three model families, with improvement ranging from 43.1% to 292.3%. Furthermore, we demonstrate that optimizing this consistency measure can improve explanation faithfulness by up to 9.7%.
title A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.19299