Selective Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Paes, Lucas Monteiro, Wei, Dennis, Calmon, Flavio P.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910464232914944
author Paes, Lucas Monteiro
Wei, Dennis
Calmon, Flavio P.
author_facet Paes, Lucas Monteiro
Wei, Dennis
Calmon, Flavio P.
contents Feature attribution methods explain black-box machine learning (ML) models by assigning importance scores to input features. These methods can be computationally expensive for large ML models. To address this challenge, there has been increasing efforts to develop amortized explainers, where a machine learning model is trained to predict feature attribution scores with only one inference. Despite their efficiency, amortized explainers can produce inaccurate predictions and misleading explanations. In this paper, we propose selective explanations, a novel feature attribution method that (i) detects when amortized explainers generate low-quality explanations and (ii) improves these explanations using a technique called explanations with initial guess. Our selective explanation method allows practitioners to specify the fraction of samples that receive explanations with initial guess, offering a principled way to bridge the gap between amortized explainers and their high-quality counterparts.
format Preprint
id arxiv_https___arxiv_org_abs_2405_19562
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Selective Explanations
Paes, Lucas Monteiro
Wei, Dennis
Calmon, Flavio P.
Computers and Society
Computation and Language
Machine Learning
Feature attribution methods explain black-box machine learning (ML) models by assigning importance scores to input features. These methods can be computationally expensive for large ML models. To address this challenge, there has been increasing efforts to develop amortized explainers, where a machine learning model is trained to predict feature attribution scores with only one inference. Despite their efficiency, amortized explainers can produce inaccurate predictions and misleading explanations. In this paper, we propose selective explanations, a novel feature attribution method that (i) detects when amortized explainers generate low-quality explanations and (ii) improves these explanations using a technique called explanations with initial guess. Our selective explanation method allows practitioners to specify the fraction of samples that receive explanations with initial guess, offering a principled way to bridge the gap between amortized explainers and their high-quality counterparts.
title Selective Explanations
topic Computers and Society
Computation and Language
Machine Learning
url https://arxiv.org/abs/2405.19562