Provable Robust Saliency-based Explanations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Chao, Guo, Chenghua, Chen, Rufeng, Ma, Guixiang, Zeng, Ming, Liao, Xiangwen, Zhang, Xi, Xie, Sihong
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913625735692288
author Chen, Chao
Guo, Chenghua
Chen, Rufeng
Ma, Guixiang
Zeng, Ming
Liao, Xiangwen
Zhang, Xi
Xie, Sihong
author_facet Chen, Chao
Guo, Chenghua
Chen, Rufeng
Ma, Guixiang
Zeng, Ming
Liao, Xiangwen
Zhang, Xi
Xie, Sihong
contents To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the $\ell_p$ distance for stability assessment, which diverges from human perception. Besides, existing adversarial training (AT) associated with intensive computations may lead to an arms race. To address these challenges, we introduce a novel metric to assess the stability of top-$k$ salient features. We introduce R2ET which trains for stable explanation by efficient and effective regularizer, and analyze R2ET by multi-objective optimization to prove numerical and statistical stability of explanations. Moreover, theoretical connections between R2ET and certified robustness justify R2ET's stability in all attacks. Extensive experiments across various data modalities and model architectures show that R2ET achieves superior stability against stealthy attacks, and generalizes effectively across different explanation methods.
format Preprint
id arxiv_https___arxiv_org_abs_2212_14106
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Provable Robust Saliency-based Explanations
Chen, Chao
Guo, Chenghua
Chen, Rufeng
Ma, Guixiang
Zeng, Ming
Liao, Xiangwen
Zhang, Xi
Xie, Sihong
Machine Learning
Artificial Intelligence
To foster trust in machine learning models, explanations must be faithful and stable for consistent insights. Existing relevant works rely on the $\ell_p$ distance for stability assessment, which diverges from human perception. Besides, existing adversarial training (AT) associated with intensive computations may lead to an arms race. To address these challenges, we introduce a novel metric to assess the stability of top-$k$ salient features. We introduce R2ET which trains for stable explanation by efficient and effective regularizer, and analyze R2ET by multi-objective optimization to prove numerical and statistical stability of explanations. Moreover, theoretical connections between R2ET and certified robustness justify R2ET's stability in all attacks. Extensive experiments across various data modalities and model architectures show that R2ET achieves superior stability against stealthy attacks, and generalizes effectively across different explanation methods.
title Provable Robust Saliency-based Explanations
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2212.14106