Neural network fragile watermarking with no model performance degradation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Yin, Zhaoxia, Yin, Heng, Zhang, Xinpeng
Format: Preprint
Published: 2022
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910731564220416
author Yin, Zhaoxia
Yin, Heng
Zhang, Xinpeng
author_facet Yin, Zhaoxia
Yin, Heng
Zhang, Xinpeng
contents Deep neural networks are vulnerable to malicious fine-tuning attacks such as data poisoning and backdoor attacks. Therefore, in recent research, it is proposed how to detect malicious fine-tuning of neural network models. However, it usually negatively affects the performance of the protected model. Thus, we propose a novel neural network fragile watermarking with no model performance degradation. In the process of watermarking, we train a generative model with the specific loss function and secret key to generate triggers that are sensitive to the fine-tuning of the target classifier. In the process of verifying, we adopt the watermarked classifier to get labels of each fragile trigger. Then, malicious fine-tuning can be detected by comparing secret keys and labels. Experiments on classic datasets and classifiers show that the proposed method can effectively detect model malicious fine-tuning with no model performance degradation.
format Preprint
id arxiv_https___arxiv_org_abs_2208_07585
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle Neural network fragile watermarking with no model performance degradation
Yin, Zhaoxia
Yin, Heng
Zhang, Xinpeng
Computer Vision and Pattern Recognition
Deep neural networks are vulnerable to malicious fine-tuning attacks such as data poisoning and backdoor attacks. Therefore, in recent research, it is proposed how to detect malicious fine-tuning of neural network models. However, it usually negatively affects the performance of the protected model. Thus, we propose a novel neural network fragile watermarking with no model performance degradation. In the process of watermarking, we train a generative model with the specific loss function and secret key to generate triggers that are sensitive to the fine-tuning of the target classifier. In the process of verifying, we adopt the watermarked classifier to get labels of each fragile trigger. Then, malicious fine-tuning can be detected by comparing secret keys and labels. Experiments on classic datasets and classifiers show that the proposed method can effectively detect model malicious fine-tuning with no model performance degradation.
title Neural network fragile watermarking with no model performance degradation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2208.07585