P2Mark: Plug-and-play Parameter-level Watermarking for Neural Speech Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Ren, Yong, Yi, Jiangyan, Wang, Tao, Tao, Jianhua, Lian, Zheng, Wen, Zhengqi, Li, Chenxing, Fu, Ruibo, Bai, Ye, Zhang, Xiaohui |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
par: Zhou, Junzuo, et autres
Publié: (2024)
par: Zhou, Junzuo, et autres
Publié: (2024)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
par: Zhou, Junzuo, et autres
Publié: (2024)
par: Zhou, Junzuo, et autres
Publié: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
par: Fan, Cunhang, et autres
Publié: (2023)
par: Fan, Cunhang, et autres
Publié: (2023)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)
par: Ren, Yong, et autres
Publié: (2026)
Residual Speaker Representation for One-Shot Voice Conversion
par: Xu, Le, et autres
Publié: (2023)
par: Xu, Le, et autres
Publié: (2023)
Towards Robust Audio Deepfake Detection: A Evolving Benchmark for Continual Learning
par: Zhang, Xiaohui, et autres
Publié: (2024)
par: Zhang, Xiaohui, et autres
Publié: (2024)
Fewer-token Neural Speech Codec with Time-invariant Codes
par: Ren, Yong, et autres
Publié: (2023)
par: Ren, Yong, et autres
Publié: (2023)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
par: Zhang, Chu Yuan, et autres
Publié: (2023)
par: Zhang, Chu Yuan, et autres
Publié: (2023)
ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
par: Gu, Hao, et autres
Publié: (2025)
par: Gu, Hao, et autres
Publié: (2025)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
par: Shi, Shuchen, et autres
Publié: (2024)
par: Shi, Shuchen, et autres
Publié: (2024)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
par: Qiang, Chunyu, et autres
Publié: (2025)
par: Qiang, Chunyu, et autres
Publié: (2025)
Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio Generation
par: Guo, Hongming, et autres
Publié: (2024)
par: Guo, Hongming, et autres
Publié: (2024)
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
par: Zeng, Siding, et autres
Publié: (2024)
par: Zeng, Siding, et autres
Publié: (2024)
VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing
par: Qiang, Chunyu, et autres
Publié: (2024)
par: Qiang, Chunyu, et autres
Publié: (2024)
DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-Speech
par: Qi, Xin, et autres
Publié: (2024)
par: Qi, Xin, et autres
Publié: (2024)
EmoFake: An Initial Dataset for Emotion Fake Audio Detection
par: Zhao, Yan, et autres
Publié: (2022)
par: Zhao, Yan, et autres
Publié: (2022)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
par: Fu, Ruibo, et autres
Publié: (2024)
par: Fu, Ruibo, et autres
Publié: (2024)
Fake Speech Wild: Detecting Deepfake Speech on Social Media Platform
par: Xie, Yuankun, et autres
Publié: (2025)
par: Xie, Yuankun, et autres
Publié: (2025)
RPRA-ADD: Forgery Trace Enhancement-Driven Audio Deepfake Detection
par: Fu, Ruibo, et autres
Publié: (2025)
par: Fu, Ruibo, et autres
Publié: (2025)
Generalized Source Tracing: Detecting Novel Audio Deepfake Algorithm with Real Emphasis and Fake Dispersion Strategy
par: Xie, Yuankun, et autres
Publié: (2024)
par: Xie, Yuankun, et autres
Publié: (2024)
Manipulated Regions Localization For Partially Deepfake Audio: A Survey
par: He, Jiayi, et autres
Publié: (2025)
par: He, Jiayi, et autres
Publié: (2025)
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
par: Xiong, Chenxu, et autres
Publié: (2024)
par: Xiong, Chenxu, et autres
Publié: (2024)
ADD 2022: the First Audio Deep Synthesis Detection Challenge
par: Yi, Jiangyan, et autres
Publié: (2022)
par: Yi, Jiangyan, et autres
Publié: (2022)
Audio Deepfake Attribution: An Initial Dataset and Investigation
par: Yan, Xinrui, et autres
Publié: (2022)
par: Yan, Xinrui, et autres
Publié: (2022)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
par: Wang, Zhiyong, et autres
Publié: (2024)
par: Wang, Zhiyong, et autres
Publié: (2024)
EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
par: Qi, Xin, et autres
Publié: (2024)
par: Qi, Xin, et autres
Publié: (2024)
Spatial Reconstructed Local Attention Res2Net with F0 Subband for Fake Speech Detection
par: Fan, Cunhang, et autres
Publié: (2023)
par: Fan, Cunhang, et autres
Publié: (2023)
Neural Codec Source Tracing: Toward Comprehensive Attribution in Open-Set Condition
par: Xie, Yuankun, et autres
Publié: (2025)
par: Xie, Yuankun, et autres
Publié: (2025)
M3-TTS: Multi-modal DiT Alignment & Mel-latent for Zero-shot High-fidelity Speech Synthesis
par: Wang, Xiaopeng, et autres
Publié: (2025)
par: Wang, Xiaopeng, et autres
Publié: (2025)
SceneFake: An Initial Dataset and Benchmarks for Scene Fake Audio Detection
par: Yi, Jiangyan, et autres
Publié: (2022)
par: Yi, Jiangyan, et autres
Publié: (2022)
Latent-Mark: An Audio Watermark Robust to Neural Resynthesis
par: Chen, Yen-Shan, et autres
Publié: (2026)
par: Chen, Yen-Shan, et autres
Publié: (2026)
WavMark: Watermarking for Audio Generation
par: Chen, Guangyu, et autres
Publié: (2023)
par: Chen, Guangyu, et autres
Publié: (2023)
Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
par: Yan, Xinrui, et autres
Publié: (2024)
par: Yan, Xinrui, et autres
Publié: (2024)
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
par: Yi, Jiangyan, et autres
Publié: (2024)
par: Yi, Jiangyan, et autres
Publié: (2024)
The Codecfake Dataset and Countermeasures for the Universally Detection of Deepfake Audio
par: Xie, Yuankun, et autres
Publié: (2024)
par: Xie, Yuankun, et autres
Publié: (2024)
Does Current Deepfake Audio Detection Model Effectively Detect ALM-based Deepfake Audio?
par: Xie, Yuankun, et autres
Publié: (2024)
par: Xie, Yuankun, et autres
Publié: (2024)
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
par: Wang, Xiaopeng, et autres
Publié: (2024)
par: Wang, Xiaopeng, et autres
Publié: (2024)
Generalized Fake Audio Detection via Deep Stable Learning
par: Wang, Zhiyong, et autres
Publié: (2024)
par: Wang, Zhiyong, et autres
Publié: (2024)
A Noval Feature via Color Quantisation for Fake Audio Detection
par: Wang, Zhiyong, et autres
Publié: (2024)
par: Wang, Zhiyong, et autres
Publié: (2024)
Documents similaires
-
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
par: Zhou, Junzuo, et autres
Publié: (2024) -
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
par: Ren, Yong, et autres
Publié: (2026) -
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
par: Zhou, Junzuo, et autres
Publié: (2024) -
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
par: Fan, Cunhang, et autres
Publié: (2023) -
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
par: Ren, Yong, et autres
Publié: (2026)