Unlearning Inversion Attacks for Graph Neural Networks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiahao, Wang, Yilong, Zhang, Zhiwei, Liu, Xiaorui, Wang, Suhang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908695263182848
author Zhang, Jiahao
Wang, Yilong
Zhang, Zhiwei
Liu, Xiaorui
Wang, Suhang
author_facet Zhang, Jiahao
Wang, Yilong
Zhang, Zhiwei
Liu, Xiaorui
Wang, Suhang
contents Graph unlearning methods aim to efficiently remove the impact of sensitive data from trained GNNs without full retraining, assuming that deleted information cannot be recovered. In this work, we challenge this assumption by introducing the graph unlearning inversion attack: given only black-box access to an unlearned GNN and partial graph knowledge, can an adversary reconstruct the removed edges? We identify two key challenges: varying probability-similarity thresholds for unlearned versus retained edges, and the difficulty of locating unlearned edge endpoints, and address them with TrendAttack. First, we derive and exploit the confidence pitfall, a theoretical and empirical pattern showing that nodes adjacent to unlearned edges exhibit a large drop in model confidence. Second, we design an adaptive prediction mechanism that applies different similarity thresholds to unlearned and other membership edges. Our framework flexibly integrates existing membership inference techniques and extends them with trend features. Experiments on four real-world datasets demonstrate that TrendAttack significantly outperforms state-of-the-art GNN membership inference baselines, exposing a critical privacy vulnerability in current graph unlearning methods.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00808
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlearning Inversion Attacks for Graph Neural Networks
Zhang, Jiahao
Wang, Yilong
Zhang, Zhiwei
Liu, Xiaorui
Wang, Suhang
Machine Learning
Artificial Intelligence
Cryptography and Security
Graph unlearning methods aim to efficiently remove the impact of sensitive data from trained GNNs without full retraining, assuming that deleted information cannot be recovered. In this work, we challenge this assumption by introducing the graph unlearning inversion attack: given only black-box access to an unlearned GNN and partial graph knowledge, can an adversary reconstruct the removed edges? We identify two key challenges: varying probability-similarity thresholds for unlearned versus retained edges, and the difficulty of locating unlearned edge endpoints, and address them with TrendAttack. First, we derive and exploit the confidence pitfall, a theoretical and empirical pattern showing that nodes adjacent to unlearned edges exhibit a large drop in model confidence. Second, we design an adaptive prediction mechanism that applies different similarity thresholds to unlearned and other membership edges. Our framework flexibly integrates existing membership inference techniques and extends them with trend features. Experiments on four real-world datasets demonstrate that TrendAttack significantly outperforms state-of-the-art GNN membership inference baselines, exposing a critical privacy vulnerability in current graph unlearning methods.
title Unlearning Inversion Attacks for Graph Neural Networks
topic Machine Learning
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2506.00808