Transparent Trade-offs between Properties of Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tadesse, Hiwot Belay, Hüyük, Alihan, Yacoby, Yaniv, Pan, Weiwei, Doshi-Velez, Finale
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909697712324608
author Tadesse, Hiwot Belay
Hüyük, Alihan
Yacoby, Yaniv
Pan, Weiwei
Doshi-Velez, Finale
author_facet Tadesse, Hiwot Belay
Hüyük, Alihan
Yacoby, Yaniv
Pan, Weiwei
Doshi-Velez, Finale
contents When explaining black-box machine learning models, it's often important for explanations to have certain desirable properties. Most existing methods `encourage' desirable properties in their construction of explanations. In this work, we demonstrate that these forms of encouragement do not consistently create explanations with the properties that are supposedly being targeted. Moreover, they do not allow for any control over which properties are prioritized when different properties are at odds with each other. We propose to directly optimize explanations for desired properties. Our direct approach not only produces explanations with optimal properties more consistently but also empowers users to control trade-offs between different properties, allowing them to create explanations with exactly what is needed for a particular task.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23880
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transparent Trade-offs between Properties of Explanations
Tadesse, Hiwot Belay
Hüyük, Alihan
Yacoby, Yaniv
Pan, Weiwei
Doshi-Velez, Finale
Machine Learning
When explaining black-box machine learning models, it's often important for explanations to have certain desirable properties. Most existing methods `encourage' desirable properties in their construction of explanations. In this work, we demonstrate that these forms of encouragement do not consistently create explanations with the properties that are supposedly being targeted. Moreover, they do not allow for any control over which properties are prioritized when different properties are at odds with each other. We propose to directly optimize explanations for desired properties. Our direct approach not only produces explanations with optimal properties more consistently but also empowers users to control trade-offs between different properties, allowing them to create explanations with exactly what is needed for a particular task.
title Transparent Trade-offs between Properties of Explanations
topic Machine Learning
url https://arxiv.org/abs/2410.23880