Informative Post-Hoc Explanations Only Exist for Simple Functions

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Günther, Eric, Szabados, Balázs, Bhattacharjee, Robi, Bordt, Sebastian, von Luxburg, Ulrike
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916901406375936
author Günther, Eric
Szabados, Balázs
Bhattacharjee, Robi
Bordt, Sebastian
von Luxburg, Ulrike
author_facet Günther, Eric
Szabados, Balázs
Bhattacharjee, Robi
Bordt, Sebastian
von Luxburg, Ulrike
contents Many researchers have suggested that local post-hoc explanation algorithms can be used to gain insights into the behavior of complex machine learning models. However, theoretical guarantees about such algorithms only exist for simple decision functions, and it is unclear whether and under which assumptions similar results might exist for complex models. In this paper, we introduce a general, learning-theory-based framework for what it means for an explanation to provide information about a decision function. We call an explanation informative if it serves to reduce the complexity of the space of plausible decision functions. With this approach, we show that many popular explanation algorithms are not informative when applied to complex decision functions, providing a rigorous mathematical rejection of the idea that it should be possible to explain any model. We then derive conditions under which different explanation algorithms become informative. These are often stronger than what one might expect. For example, gradient explanations and counterfactual explanations are non-informative with respect to the space of differentiable functions, and SHAP and anchor explanations are not informative with respect to the space of decision trees. Based on these results, we discuss how explanation algorithms can be modified to become informative. While the proposed analysis of explanation algorithms is mathematical, we argue that it holds strong implications for the practical applicability of these algorithms, particularly for auditing, regulation, and high-risk applications of AI.
format Preprint
id arxiv_https___arxiv_org_abs_2508_11441
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Informative Post-Hoc Explanations Only Exist for Simple Functions
Günther, Eric
Szabados, Balázs
Bhattacharjee, Robi
Bordt, Sebastian
von Luxburg, Ulrike
Machine Learning
Artificial Intelligence
Many researchers have suggested that local post-hoc explanation algorithms can be used to gain insights into the behavior of complex machine learning models. However, theoretical guarantees about such algorithms only exist for simple decision functions, and it is unclear whether and under which assumptions similar results might exist for complex models. In this paper, we introduce a general, learning-theory-based framework for what it means for an explanation to provide information about a decision function. We call an explanation informative if it serves to reduce the complexity of the space of plausible decision functions. With this approach, we show that many popular explanation algorithms are not informative when applied to complex decision functions, providing a rigorous mathematical rejection of the idea that it should be possible to explain any model. We then derive conditions under which different explanation algorithms become informative. These are often stronger than what one might expect. For example, gradient explanations and counterfactual explanations are non-informative with respect to the space of differentiable functions, and SHAP and anchor explanations are not informative with respect to the space of decision trees. Based on these results, we discuss how explanation algorithms can be modified to become informative. While the proposed analysis of explanation algorithms is mathematical, we argue that it holds strong implications for the practical applicability of these algorithms, particularly for auditing, regulation, and high-risk applications of AI.
title Informative Post-Hoc Explanations Only Exist for Simple Functions
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2508.11441