Explainable Graph Neural Networks Under Fire

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Li, Zhong, Geisler, Simon, Wang, Yuhang, Günnemann, Stephan, van Leeuwen, Matthijs
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866910655415582720
author Li, Zhong
Geisler, Simon
Wang, Yuhang
Günnemann, Stephan
van Leeuwen, Matthijs
author_facet Li, Zhong
Geisler, Simon
Wang, Yuhang
Günnemann, Stephan
van Leeuwen, Matthijs
contents Predictions made by graph neural networks (GNNs) usually lack interpretability due to their complex computational behavior and the abstract nature of graphs. In an attempt to tackle this, many GNN explanation methods have emerged. Their goal is to explain a model's predictions and thereby obtain trust when GNN models are deployed in decision critical applications. Most GNN explanation methods work in a post-hoc manner and provide explanations in the form of a small subset of important edges and/or nodes. In this paper we demonstrate that these explanations can unfortunately not be trusted, as common GNN explanation methods turn out to be highly susceptible to adversarial perturbations. That is, even small perturbations of the original graph structure that preserve the model's predictions may yield drastically different explanations. This calls into question the trustworthiness and practical utility of post-hoc explanation methods for GNNs. To be able to attack GNN explanation models, we devise a novel attack method dubbed \textit{GXAttack}, the first \textit{optimization-based} adversarial white-box attack method for post-hoc GNN explanations under such settings. Due to the devastating effectiveness of our attack, we call for an adversarial evaluation of future GNN explainers to demonstrate their robustness. For reproducibility, our code is available via GitHub.
format Preprint
id arxiv_https___arxiv_org_abs_2406_06417
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Explainable Graph Neural Networks Under Fire
Li, Zhong
Geisler, Simon
Wang, Yuhang
Günnemann, Stephan
van Leeuwen, Matthijs
Machine Learning
Artificial Intelligence
Predictions made by graph neural networks (GNNs) usually lack interpretability due to their complex computational behavior and the abstract nature of graphs. In an attempt to tackle this, many GNN explanation methods have emerged. Their goal is to explain a model's predictions and thereby obtain trust when GNN models are deployed in decision critical applications. Most GNN explanation methods work in a post-hoc manner and provide explanations in the form of a small subset of important edges and/or nodes. In this paper we demonstrate that these explanations can unfortunately not be trusted, as common GNN explanation methods turn out to be highly susceptible to adversarial perturbations. That is, even small perturbations of the original graph structure that preserve the model's predictions may yield drastically different explanations. This calls into question the trustworthiness and practical utility of post-hoc explanation methods for GNNs. To be able to attack GNN explanation models, we devise a novel attack method dubbed \textit{GXAttack}, the first \textit{optimization-based} adversarial white-box attack method for post-hoc GNN explanations under such settings. Due to the devastating effectiveness of our attack, we call for an adversarial evaluation of future GNN explainers to demonstrate their robustness. For reproducibility, our code is available via GitHub.
title Explainable Graph Neural Networks Under Fire
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2406.06417