Interpret the Predictions of Deep Networks via Re-Label Distillation

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Hua, Yingying, Ge, Shiming, Zhang, Daichi
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916402536906752
author Hua, Yingying
Ge, Shiming
Zhang, Daichi
author_facet Hua, Yingying
Ge, Shiming
Zhang, Daichi
contents Interpreting the predictions of a black-box deep network can facilitate the reliability of its deployment. In this work, we propose a re-label distillation approach to learn a direct map from the input to the prediction in a self-supervision manner. The image is projected into a VAE subspace to generate some synthetic images by randomly perturbing its latent vector. Then, these synthetic images can be annotated into one of two classes by identifying whether their labels shift. After that, using the labels annotated by the deep network as teacher, a linear student model is trained to approximate the annotations by mapping these synthetic images to the classes. In this manner, these re-labeled synthetic images can well describe the local classification mechanism of the deep network, and the learned student can provide a more intuitive explanation towards the predictions. Extensive experiments verify the effectiveness of our approach qualitatively and quantitatively.
format Preprint
id arxiv_https___arxiv_org_abs_2409_13137
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Interpret the Predictions of Deep Networks via Re-Label Distillation
Hua, Yingying
Ge, Shiming
Zhang, Daichi
Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
Interpreting the predictions of a black-box deep network can facilitate the reliability of its deployment. In this work, we propose a re-label distillation approach to learn a direct map from the input to the prediction in a self-supervision manner. The image is projected into a VAE subspace to generate some synthetic images by randomly perturbing its latent vector. Then, these synthetic images can be annotated into one of two classes by identifying whether their labels shift. After that, using the labels annotated by the deep network as teacher, a linear student model is trained to approximate the annotations by mapping these synthetic images to the classes. In this manner, these re-labeled synthetic images can well describe the local classification mechanism of the deep network, and the learned student can provide a more intuitive explanation towards the predictions. Extensive experiments verify the effectiveness of our approach qualitatively and quantitatively.
title Interpret the Predictions of Deep Networks via Re-Label Distillation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2409.13137