Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiale, Rao, Bosen, Zhu, Chengcheng, Sun, Xiaobing, Li, Qingming, Hu, Haibo, Luo, Xiapu, Ye, Qingqing, Ji, Shouling
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909453605928960
author Zhang, Jiale
Rao, Bosen
Zhu, Chengcheng
Sun, Xiaobing
Li, Qingming
Hu, Haibo
Luo, Xiapu
Ye, Qingqing
Ji, Shouling
author_facet Zhang, Jiale
Rao, Bosen
Zhu, Chengcheng
Sun, Xiaobing
Li, Qingming
Hu, Haibo
Luo, Xiapu
Ye, Qingqing
Ji, Shouling
contents Graph Neural Networks (GNNs) have achieved remarkable performance through their message-passing mechanism. However, recent studies have highlighted the vulnerability of GNNs to backdoor attacks, which can lead the model to misclassify graphs with attached triggers as the target class. The effectiveness of recent promising defense techniques, such as fine-tuning or distillation, is heavily contingent on having comprehensive knowledge of the sufficient training dataset. Empirical studies have shown that fine-tuning methods require a clean dataset of 20% to reduce attack accuracy to below 25%, while distillation methods require a clean dataset of 15%. However, obtaining such a large amount of clean data is commonly impractical. In this paper, we propose a practical backdoor mitigation framework, denoted as GRAPHNAD, which can capture high-quality intermediate-layer representations in GNNs to enhance the distillation process with limited clean data. To achieve this, we address the following key questions: How to identify the appropriate attention representations in graphs for distillation? How to enhance distillation with limited data? By adopting the graph attention transfer method, GRAPHNAD can effectively align the intermediate-layer attention representations of the backdoored model with that of the teacher model, forcing the backdoor neurons to transform into benign ones. Besides, we extract the relation maps from intermediate-layer transformation and enforce the relation maps of the backdoored model to be consistent with that of the teacher model, thereby ensuring model accuracy while further reducing the influence of backdoors. Extensive experimental results show that by fine-tuning a teacher model with only 3% of the clean data, GRAPHNAD can reduce the attack success rate to below 5%.
format Preprint
id arxiv_https___arxiv_org_abs_2501_05835
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
Zhang, Jiale
Rao, Bosen
Zhu, Chengcheng
Sun, Xiaobing
Li, Qingming
Hu, Haibo
Luo, Xiapu
Ye, Qingqing
Ji, Shouling
Machine Learning
Cryptography and Security
Graph Neural Networks (GNNs) have achieved remarkable performance through their message-passing mechanism. However, recent studies have highlighted the vulnerability of GNNs to backdoor attacks, which can lead the model to misclassify graphs with attached triggers as the target class. The effectiveness of recent promising defense techniques, such as fine-tuning or distillation, is heavily contingent on having comprehensive knowledge of the sufficient training dataset. Empirical studies have shown that fine-tuning methods require a clean dataset of 20% to reduce attack accuracy to below 25%, while distillation methods require a clean dataset of 15%. However, obtaining such a large amount of clean data is commonly impractical. In this paper, we propose a practical backdoor mitigation framework, denoted as GRAPHNAD, which can capture high-quality intermediate-layer representations in GNNs to enhance the distillation process with limited clean data. To achieve this, we address the following key questions: How to identify the appropriate attention representations in graphs for distillation? How to enhance distillation with limited data? By adopting the graph attention transfer method, GRAPHNAD can effectively align the intermediate-layer attention representations of the backdoored model with that of the teacher model, forcing the backdoor neurons to transform into benign ones. Besides, we extract the relation maps from intermediate-layer transformation and enforce the relation maps of the backdoored model to be consistent with that of the teacher model, thereby ensuring model accuracy while further reducing the influence of backdoors. Extensive experimental results show that by fine-tuning a teacher model with only 3% of the clean data, GRAPHNAD can reduce the attack success rate to below 5%.
title Fine-tuning is Not Fine: Mitigating Backdoor Attacks in GNNs with Limited Clean Data
topic Machine Learning
Cryptography and Security
url https://arxiv.org/abs/2501.05835