Speech Watermarking with Discrete Intermediate Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ji, Shengpeng, Jiang, Ziyue, Zuo, Jialong, Fang, Minghui, Chen, Yifu, Jin, Tao, Zhao, Zhou
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912160758628352
author Ji, Shengpeng
Jiang, Ziyue
Zuo, Jialong
Fang, Minghui
Chen, Yifu
Jin, Tao
Zhao, Zhou
author_facet Ji, Shengpeng
Jiang, Ziyue
Zuo, Jialong
Fang, Minghui
Chen, Yifu
Jin, Tao
Zhao, Zhou
contents Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity. Audio samples are available at https://DiscreteWM.github.io/discrete_wm.
format Preprint
id arxiv_https___arxiv_org_abs_2412_13917
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Speech Watermarking with Discrete Intermediate Representations
Ji, Shengpeng
Jiang, Ziyue
Zuo, Jialong
Fang, Minghui
Chen, Yifu
Jin, Tao
Zhao, Zhou
Audio and Speech Processing
Machine Learning
Sound
Signal Processing
Speech watermarking techniques can proactively mitigate the potential harmful consequences of instant voice cloning techniques. These techniques involve the insertion of signals into speech that are imperceptible to humans but can be detected by algorithms. Previous approaches typically embed watermark messages into continuous space. However, intuitively, embedding watermark information into robust discrete latent space can significantly improve the robustness of watermarking systems. In this paper, we propose DiscreteWM, a novel speech watermarking framework that injects watermarks into the discrete intermediate representations of speech. Specifically, we map speech into discrete latent space with a vector-quantized autoencoder and inject watermarks by changing the modular arithmetic relation of discrete IDs. To ensure the imperceptibility of watermarks, we also propose a manipulator model to select the candidate tokens for watermark embedding. Experimental results demonstrate that our framework achieves state-of-the-art performance in robustness and imperceptibility, simultaneously. Moreover, our flexible frame-wise approach can serve as an efficient solution for both voice cloning detection and information hiding. Additionally, DiscreteWM can encode 1 to 150 bits of watermark information within a 1-second speech clip, indicating its encoding capacity. Audio samples are available at https://DiscreteWM.github.io/discrete_wm.
title Speech Watermarking with Discrete Intermediate Representations
topic Audio and Speech Processing
Machine Learning
Sound
Signal Processing
url https://arxiv.org/abs/2412.13917