GNN Explanations that do not Explain and How to find Them

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Azzolin, Steve, Teso, Stefano, Lepri, Bruno, Passerini, Andrea, Malhotra, Sagar
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908858217136128
author Azzolin, Steve
Teso, Stefano
Lepri, Bruno
Passerini, Andrea
Malhotra, Sagar
author_facet Azzolin, Steve
Teso, Stefano
Lepri, Bruno
Passerini, Andrea
Malhotra, Sagar
contents Explanations provided by Self-explainable Graph Neural Networks (SE-GNNs) are fundamental for understanding the model's inner workings and for identifying potential misuse of sensitive attributes. Although recent works have highlighted that these explanations can be suboptimal and potentially misleading, a characterization of their failure cases is unavailable. In this work, we identify a critical failure of SE-GNN explanations: explanations can be unambiguously unrelated to how the SE-GNNs infer labels. We show that, on the one hand, many SE-GNNs can achieve optimal true risk while producing these degenerate explanations, and on the other, most faithfulness metrics can fail to identify these failure modes. Our empirical analysis reveals that degenerate explanations can be maliciously planted (allowing an attacker to hide the use of sensitive attributes) and can also emerge naturally, highlighting the need for reliable auditing. To address this, we introduce a novel faithfulness metric that reliably marks degenerate explanations as unfaithful, in both malicious and natural settings. Our code is available at https://github.com/steveazzolin/gnn_deg_expl.
format Preprint
id arxiv_https___arxiv_org_abs_2601_20815
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GNN Explanations that do not Explain and How to find Them
Azzolin, Steve
Teso, Stefano
Lepri, Bruno
Passerini, Andrea
Malhotra, Sagar
Machine Learning
Artificial Intelligence
Explanations provided by Self-explainable Graph Neural Networks (SE-GNNs) are fundamental for understanding the model's inner workings and for identifying potential misuse of sensitive attributes. Although recent works have highlighted that these explanations can be suboptimal and potentially misleading, a characterization of their failure cases is unavailable. In this work, we identify a critical failure of SE-GNN explanations: explanations can be unambiguously unrelated to how the SE-GNNs infer labels. We show that, on the one hand, many SE-GNNs can achieve optimal true risk while producing these degenerate explanations, and on the other, most faithfulness metrics can fail to identify these failure modes. Our empirical analysis reveals that degenerate explanations can be maliciously planted (allowing an attacker to hide the use of sensitive attributes) and can also emerge naturally, highlighting the need for reliable auditing. To address this, we introduce a novel faithfulness metric that reliably marks degenerate explanations as unfaithful, in both malicious and natural settings. Our code is available at https://github.com/steveazzolin/gnn_deg_expl.
title GNN Explanations that do not Explain and How to find Them
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2601.20815