Forging the Unforgeable: On the Feasibility of Counterfeit Watermarks in Backdoor-Based Dataset Ownership Verification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Zhiying, Liu, Zhi, Liu, Dongjie, Zhuo, Shengda, Geng, Guanggang, Fan, Zhaoxin, Lyu, Shanxiang, Jin, Xiaobo, Weng, Jian
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915852475957248
author Li, Zhiying
Liu, Zhi
Liu, Dongjie
Zhuo, Shengda
Geng, Guanggang
Fan, Zhaoxin
Lyu, Shanxiang
Jin, Xiaobo
Weng, Jian
author_facet Li, Zhiying
Liu, Zhi
Liu, Dongjie
Zhuo, Shengda
Geng, Guanggang
Fan, Zhaoxin
Lyu, Shanxiang
Jin, Xiaobo
Weng, Jian
contents Backdoor watermarking has emerged as the predominant approach for protecting public datasets, enabling dataset ownership verification (DOV) through embedded triggers that induce predefined model behaviors. While existing works assume that DOV results can serve as reliable evidence for copyright infringement claims, we argue that this assumption is fundamentally flawed. In this paper, we expose critical vulnerabilities in current backdoor watermarking schemes by demonstrating that attackers can forge watermarks that are statistically indistinguishable from the original ones, thereby evading infringement allegations. Specifically, we propose a Forged Watermark Generator (FW-Gen), a lightweight variational autoencoder-based framework that generates forged watermarks preserving the statistical properties of original watermarks while exhibiting distinct visual patterns. Our attack operates under a realistic threat model where an accused attacker, upon receiving an infringement claim, extracts watermark information from the protected dataset and produces counterfeit evidence to refute the allegation. Extensive experiments across six backdoor watermarking methods, two benchmark datasets, and two model architectures demonstrate that forged watermarks achieve equivalent or superior statistical significance in hypothesis testing compared to original watermarks. These findings reveal that current DOV mechanisms are insufficient as standalone evidence for copyright disputes and call for more robust dataset protection schemes.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15450
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Forging the Unforgeable: On the Feasibility of Counterfeit Watermarks in Backdoor-Based Dataset Ownership Verification
Li, Zhiying
Liu, Zhi
Liu, Dongjie
Zhuo, Shengda
Geng, Guanggang
Fan, Zhaoxin
Lyu, Shanxiang
Jin, Xiaobo
Weng, Jian
Cryptography and Security
Backdoor watermarking has emerged as the predominant approach for protecting public datasets, enabling dataset ownership verification (DOV) through embedded triggers that induce predefined model behaviors. While existing works assume that DOV results can serve as reliable evidence for copyright infringement claims, we argue that this assumption is fundamentally flawed. In this paper, we expose critical vulnerabilities in current backdoor watermarking schemes by demonstrating that attackers can forge watermarks that are statistically indistinguishable from the original ones, thereby evading infringement allegations. Specifically, we propose a Forged Watermark Generator (FW-Gen), a lightweight variational autoencoder-based framework that generates forged watermarks preserving the statistical properties of original watermarks while exhibiting distinct visual patterns. Our attack operates under a realistic threat model where an accused attacker, upon receiving an infringement claim, extracts watermark information from the protected dataset and produces counterfeit evidence to refute the allegation. Extensive experiments across six backdoor watermarking methods, two benchmark datasets, and two model architectures demonstrate that forged watermarks achieve equivalent or superior statistical significance in hypothesis testing compared to original watermarks. These findings reveal that current DOV mechanisms are insufficient as standalone evidence for copyright disputes and call for more robust dataset protection schemes.
title Forging the Unforgeable: On the Feasibility of Counterfeit Watermarks in Backdoor-Based Dataset Ownership Verification
topic Cryptography and Security
url https://arxiv.org/abs/2411.15450