CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Yiming, Yan, Kaiying, Shao, Shuo, Zhai, Tongqing, Xia, Shu-Tao, Qin, Zhan, Tao, Dacheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909566308974592
author Li, Yiming
Yan, Kaiying
Shao, Shuo
Zhai, Tongqing
Xia, Shu-Tao
Qin, Zhan
Tao, Dacheng
author_facet Li, Yiming
Yan, Kaiying
Shao, Shuo
Zhai, Tongqing
Xia, Shu-Tao
Qin, Zhan
Tao, Dacheng
contents With the increasing adoption of deep learning in speaker verification, large-scale speech datasets have become valuable intellectual property. To audit and prevent the unauthorized usage of these valuable released datasets, especially in commercial or open-source scenarios, we propose a novel dataset ownership verification method. Our approach introduces a clustering-based backdoor watermark (CBW), enabling dataset owners to determine whether a suspicious third-party model has been trained on a protected dataset under a black-box setting. The CBW method consists of two key stages: dataset watermarking and ownership verification. During watermarking, we implant multiple trigger patterns in the dataset to make similar samples (measured by their feature similarities) close to the same trigger while dissimilar samples are near different triggers. This ensures that any model trained on the watermarked dataset exhibits specific misclassification behaviors when exposed to trigger-embedded inputs. To verify dataset ownership, we design a hypothesis-test-based framework that statistically evaluates whether a suspicious model exhibits the expected backdoor behavior. We conduct extensive experiments on benchmark datasets, verifying the effectiveness and robustness of our method against potential adaptive attacks. The code for reproducing main experiments is available at https://github.com/Radiant0726/CBW
format Preprint
id arxiv_https___arxiv_org_abs_2503_05794
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking
Li, Yiming
Yan, Kaiying
Shao, Shuo
Zhai, Tongqing
Xia, Shu-Tao
Qin, Zhan
Tao, Dacheng
Cryptography and Security
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
With the increasing adoption of deep learning in speaker verification, large-scale speech datasets have become valuable intellectual property. To audit and prevent the unauthorized usage of these valuable released datasets, especially in commercial or open-source scenarios, we propose a novel dataset ownership verification method. Our approach introduces a clustering-based backdoor watermark (CBW), enabling dataset owners to determine whether a suspicious third-party model has been trained on a protected dataset under a black-box setting. The CBW method consists of two key stages: dataset watermarking and ownership verification. During watermarking, we implant multiple trigger patterns in the dataset to make similar samples (measured by their feature similarities) close to the same trigger while dissimilar samples are near different triggers. This ensures that any model trained on the watermarked dataset exhibits specific misclassification behaviors when exposed to trigger-embedded inputs. To verify dataset ownership, we design a hypothesis-test-based framework that statistically evaluates whether a suspicious model exhibits the expected backdoor behavior. We conduct extensive experiments on benchmark datasets, verifying the effectiveness and robustness of our method against potential adaptive attacks. The code for reproducing main experiments is available at https://github.com/Radiant0726/CBW
title CBW: Towards Dataset Ownership Verification for Speaker Verification via Clustering-based Backdoor Watermarking
topic Cryptography and Security
Artificial Intelligence
Machine Learning
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2503.05794