Extracting Explanations, Justification, and Uncertainty from Black-Box Deep Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ardis, Paul, Flenner, Arjuna
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911796203356160
author Ardis, Paul
Flenner, Arjuna
author_facet Ardis, Paul
Flenner, Arjuna
contents Deep Neural Networks (DNNs) do not inherently compute or exhibit empirically-justified task confidence. In mission critical applications, it is important to both understand associated DNN reasoning and its supporting evidence. In this paper, we propose a novel Bayesian approach to extract explanations, justifications, and uncertainty estimates from DNNs. Our approach is efficient both in terms of memory and computation, and can be applied to any black box DNN without any retraining, including applications to anomaly detection and out-of-distribution detection tasks. We validate our approach on the CIFAR-10 dataset, and show that it can significantly improve the interpretability and reliability of DNNs.
format Preprint
id arxiv_https___arxiv_org_abs_2403_08652
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Extracting Explanations, Justification, and Uncertainty from Black-Box Deep Neural Networks
Ardis, Paul
Flenner, Arjuna
Machine Learning
I.2.10
Deep Neural Networks (DNNs) do not inherently compute or exhibit empirically-justified task confidence. In mission critical applications, it is important to both understand associated DNN reasoning and its supporting evidence. In this paper, we propose a novel Bayesian approach to extract explanations, justifications, and uncertainty estimates from DNNs. Our approach is efficient both in terms of memory and computation, and can be applied to any black box DNN without any retraining, including applications to anomaly detection and out-of-distribution detection tasks. We validate our approach on the CIFAR-10 dataset, and show that it can significantly improve the interpretability and reliability of DNNs.
title Extracting Explanations, Justification, and Uncertainty from Black-Box Deep Neural Networks
topic Machine Learning
I.2.10
url https://arxiv.org/abs/2403.08652