Sufficient and Necessary Explanations (and What Lies in Between)

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bharti, Beepul, Yi, Paul, Sulam, Jeremias
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913546774773760
author Bharti, Beepul
Yi, Paul
Sulam, Jeremias
author_facet Bharti, Beepul
Yi, Paul
Sulam, Jeremias
contents As complex machine learning models continue to find applications in high-stakes decision-making scenarios, it is crucial that we can explain and understand their predictions. Post-hoc explanation methods provide useful insights by identifying important features in an input $\mathbf{x}$ with respect to the model output $f(\mathbf{x})$. In this work, we formalize and study two precise notions of feature importance for general machine learning models: sufficiency and necessity. We demonstrate how these two types of explanations, albeit intuitive and simple, can fall short in providing a complete picture of which features a model finds important. To this end, we propose a unified notion of importance that circumvents these limitations by exploring a continuum along a necessity-sufficiency axis. Our unified notion, we show, has strong ties to other popular definitions of feature importance, like those based on conditional independence and game-theoretic quantities like Shapley values. Crucially, we demonstrate how a unified perspective allows us to detect important features that could be missed by either of the previous approaches alone.
format Preprint
id arxiv_https___arxiv_org_abs_2409_20427
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sufficient and Necessary Explanations (and What Lies in Between)
Bharti, Beepul
Yi, Paul
Sulam, Jeremias
Machine Learning
Artificial Intelligence
As complex machine learning models continue to find applications in high-stakes decision-making scenarios, it is crucial that we can explain and understand their predictions. Post-hoc explanation methods provide useful insights by identifying important features in an input $\mathbf{x}$ with respect to the model output $f(\mathbf{x})$. In this work, we formalize and study two precise notions of feature importance for general machine learning models: sufficiency and necessity. We demonstrate how these two types of explanations, albeit intuitive and simple, can fall short in providing a complete picture of which features a model finds important. To this end, we propose a unified notion of importance that circumvents these limitations by exploring a continuum along a necessity-sufficiency axis. Our unified notion, we show, has strong ties to other popular definitions of feature importance, like those based on conditional independence and game-theoretic quantities like Shapley values. Crucially, we demonstrate how a unified perspective allows us to detect important features that could be missed by either of the previous approaches alone.
title Sufficient and Necessary Explanations (and What Lies in Between)
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2409.20427