Watermarking Generative Tabular Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Hengzhi, Yu, Peiyu, Ren, Junpeng, Wu, Ying Nian, Cheng, Guang
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911886699659264
author He, Hengzhi
Yu, Peiyu
Ren, Junpeng
Wu, Ying Nian
Cheng, Guang
author_facet He, Hengzhi
Yu, Peiyu
Ren, Junpeng
Wu, Ying Nian
Cheng, Guang
contents In this paper, we introduce a simple yet effective tabular data watermarking mechanism with statistical guarantees. We show theoretically that the proposed watermark can be effectively detected, while faithfully preserving the data fidelity, and also demonstrates appealing robustness against additive noise attack. The general idea is to achieve the watermarking through a strategic embedding based on simple data binning. Specifically, it divides the feature's value range into finely segmented intervals and embeds watermarks into selected ``green list" intervals. To detect the watermarks, we develop a principled statistical hypothesis-testing framework with minimal assumptions: it remains valid as long as the underlying data distribution has a continuous density function. The watermarking efficacy is demonstrated through rigorous theoretical analysis and empirical validation, highlighting its utility in enhancing the security of synthetic and real-world datasets.
format Preprint
id arxiv_https___arxiv_org_abs_2405_14018
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Watermarking Generative Tabular Data
He, Hengzhi
Yu, Peiyu
Ren, Junpeng
Wu, Ying Nian
Cheng, Guang
Cryptography and Security
Machine Learning
Applications
In this paper, we introduce a simple yet effective tabular data watermarking mechanism with statistical guarantees. We show theoretically that the proposed watermark can be effectively detected, while faithfully preserving the data fidelity, and also demonstrates appealing robustness against additive noise attack. The general idea is to achieve the watermarking through a strategic embedding based on simple data binning. Specifically, it divides the feature's value range into finely segmented intervals and embeds watermarks into selected ``green list" intervals. To detect the watermarks, we develop a principled statistical hypothesis-testing framework with minimal assumptions: it remains valid as long as the underlying data distribution has a continuous density function. The watermarking efficacy is demonstrated through rigorous theoretical analysis and empirical validation, highlighting its utility in enhancing the security of synthetic and real-world datasets.
title Watermarking Generative Tabular Data
topic Cryptography and Security
Machine Learning
Applications
url https://arxiv.org/abs/2405.14018