Interpretability Guarantees with Merlin-Arthur Classifiers

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wäldchen, Stephan, Sharma, Kartikey, Turan, Berkant, Zimmer, Max, Pokutta, Sebastian
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910378129096704
author Wäldchen, Stephan
Sharma, Kartikey
Turan, Berkant
Zimmer, Max
Pokutta, Sebastian
author_facet Wäldchen, Stephan
Sharma, Kartikey
Turan, Berkant
Zimmer, Max
Pokutta, Sebastian
contents We propose an interactive multi-agent classifier that provides provable interpretability guarantees even for complex agents such as neural networks. These guarantees consist of lower bounds on the mutual information between selected features and the classification decision. Our results are inspired by the Merlin-Arthur protocol from Interactive Proof Systems and express these bounds in terms of measurable metrics such as soundness and completeness. Compared to existing interactive setups, we rely neither on optimal agents nor on the assumption that features are distributed independently. Instead, we use the relative strength of the agents as well as the new concept of Asymmetric Feature Correlation which captures the precise kind of correlations that make interpretability guarantees difficult. We evaluate our results on two small-scale datasets where high mutual information can be verified explicitly.
format Preprint
id arxiv_https___arxiv_org_abs_2206_00759
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Interpretability Guarantees with Merlin-Arthur Classifiers
Wäldchen, Stephan
Sharma, Kartikey
Turan, Berkant
Zimmer, Max
Pokutta, Sebastian
Machine Learning
Artificial Intelligence
68T01, 91A06
I.2.0
We propose an interactive multi-agent classifier that provides provable interpretability guarantees even for complex agents such as neural networks. These guarantees consist of lower bounds on the mutual information between selected features and the classification decision. Our results are inspired by the Merlin-Arthur protocol from Interactive Proof Systems and express these bounds in terms of measurable metrics such as soundness and completeness. Compared to existing interactive setups, we rely neither on optimal agents nor on the assumption that features are distributed independently. Instead, we use the relative strength of the agents as well as the new concept of Asymmetric Feature Correlation which captures the precise kind of correlations that make interpretability guarantees difficult. We evaluate our results on two small-scale datasets where high mutual information can be verified explicitly.
title Interpretability Guarantees with Merlin-Arthur Classifiers
topic Machine Learning
Artificial Intelligence
68T01, 91A06
I.2.0
url https://arxiv.org/abs/2206.00759