Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yichuan, Li, Chengxin, Gu, Yujie
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912780603359232
author Zhang, Yichuan
Li, Chengxin
Gu, Yujie
author_facet Zhang, Yichuan
Li, Chengxin
Gu, Yujie
contents Text-to-Speech (TTS) diffusion models generate high-quality speech, which raises challenges for the model intellectual property protection and speech tracing for legal use. Audio watermarking is a promising solution. However, due to the structural differences among various TTS diffusion models, existing watermarking methods are often designed for a specific model and degrade audio quality, which limits their practical applicability. To address this dilemma, this paper proposes a universal watermarking scheme for TTS diffusion models, termed Smark. This is achieved by designing a lightweight watermark embedding framework that operates in the common reverse diffusion paradigm shared by all TTS diffusion models. To mitigate the impact on audio quality, Smark utilizes the discrete wavelet transform (DWT) to embed watermarks into the relatively stable low-frequency regions of the audio, which ensures seamless watermark-audio integration and is resistant to removal during the reverse diffusion process. Extensive experiments are conducted to evaluate the audio quality and watermark performance in various simulated real-world attack scenarios. The experimental results show that Smark achieves superior performance in both audio quality and watermark extraction accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2512_18791
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
Zhang, Yichuan
Li, Chengxin
Gu, Yujie
Sound
Artificial Intelligence
Cryptography and Security
Text-to-Speech (TTS) diffusion models generate high-quality speech, which raises challenges for the model intellectual property protection and speech tracing for legal use. Audio watermarking is a promising solution. However, due to the structural differences among various TTS diffusion models, existing watermarking methods are often designed for a specific model and degrade audio quality, which limits their practical applicability. To address this dilemma, this paper proposes a universal watermarking scheme for TTS diffusion models, termed Smark. This is achieved by designing a lightweight watermark embedding framework that operates in the common reverse diffusion paradigm shared by all TTS diffusion models. To mitigate the impact on audio quality, Smark utilizes the discrete wavelet transform (DWT) to embed watermarks into the relatively stable low-frequency regions of the audio, which ensures seamless watermark-audio integration and is resistant to removal during the reverse diffusion process. Extensive experiments are conducted to evaluate the audio quality and watermark performance in various simulated real-world attack scenarios. The experimental results show that Smark achieves superior performance in both audio quality and watermark extraction accuracy.
title Smark: A Watermark for Text-to-Speech Diffusion Models via Discrete Wavelet Transform
topic Sound
Artificial Intelligence
Cryptography and Security
url https://arxiv.org/abs/2512.18791