Explanations are a Means to an End: Decision Theoretic Explanation Evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Guo, Ziyang, Ustun, Berk, Hullman, Jessica
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911461736972288
author Guo, Ziyang
Ustun, Berk
Hullman, Jessica
author_facet Guo, Ziyang
Ustun, Berk
Hullman, Jessica
contents Explanations of model behavior are commonly evaluated via proxy properties weakly tied to the purposes explanations serve in practice. We contribute a decision theoretic framework that treats explanations as information signals valued by the expected improvement they enable on a specified decision task. This approach yields three distinct estimands: 1) a theoretical benchmark that upperbounds achievable performance by any agent with the explanation, 2) a human-complementary value that quantifies the theoretically attainable value that is not already captured by a baseline human decision policy, and 3) a behavioral value representing the causal effect of providing the explanation to human decision-makers. We instantiate these definitions in a practical validation workflow, and apply them to assess explanation potential and interpret behavioral effects in human-AI decision support and mechanistic interpretability.
format Preprint
id arxiv_https___arxiv_org_abs_2506_22740
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
Guo, Ziyang
Ustun, Berk
Hullman, Jessica
Artificial Intelligence
Machine Learning
Explanations of model behavior are commonly evaluated via proxy properties weakly tied to the purposes explanations serve in practice. We contribute a decision theoretic framework that treats explanations as information signals valued by the expected improvement they enable on a specified decision task. This approach yields three distinct estimands: 1) a theoretical benchmark that upperbounds achievable performance by any agent with the explanation, 2) a human-complementary value that quantifies the theoretically attainable value that is not already captured by a baseline human decision policy, and 3) a behavioral value representing the causal effect of providing the explanation to human decision-makers. We instantiate these definitions in a practical validation workflow, and apply them to assess explanation potential and interpret behavioral effects in human-AI decision support and mechanistic interpretability.
title Explanations are a Means to an End: Decision Theoretic Explanation Evaluation
topic Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2506.22740