WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Junzuo, Yi, Jiangyan, Ren, Yong, Tao, Jianhua, Wang, Tao, Zhang, Chu Yuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023)
von: Ren, Yong, et al.
Veröffentlicht: (2023)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
Towards Robust Audio Deepfake Detection: A Evolving Benchmark for Continual Learning
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
Residual Speaker Representation for One-Shot Voice Conversion
von: Xu, Le, et al.
Veröffentlicht: (2023)
von: Xu, Le, et al.
Veröffentlicht: (2023)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
SecoustiCodec: Cross-Modal Aligned Streaming Single-Codecbook Speech Codec
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
von: Qiang, Chunyu, et al.
Veröffentlicht: (2025)
Spatial Reconstructed Local Attention Res2Net with F0 Subband for Fake Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
von: Chen, Junyang, et al.
Veröffentlicht: (2026)
SpatialCodec: Neural Spatial Speech Coding
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
von: Xu, Zhongweiyang, et al.
Veröffentlicht: (2023)
EmoFake: An Initial Dataset for Emotion Fake Audio Detection
von: Zhao, Yan, et al.
Veröffentlicht: (2022)
von: Zhao, Yan, et al.
Veröffentlicht: (2022)
CodecFake+: A Large-Scale Neural Audio Codec-Based Deepfake Speech Dataset
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2025)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
Personalized Neural Speech Codec
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
von: Jang, Inseon, et al.
Veröffentlicht: (2024)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
von: Chen, Wenxi, et al.
Veröffentlicht: (2025)
On Improving Error Resilience of Neural End-to-End Speech Coders
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
von: Gupta, Kishan, et al.
Veröffentlicht: (2024)
A Neural Speech Codec for Noise Robust Speech Coding
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
von: Huang, Jiayi, et al.
Veröffentlicht: (2023)
BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
von: Xin, Detai, et al.
Veröffentlicht: (2024)
von: Xin, Detai, et al.
Veröffentlicht: (2024)
ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
FLY-TTS: Fast, Lightweight and High-Quality End-to-End Text-to-Speech Synthesis
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
von: Guo, Yinlin, et al.
Veröffentlicht: (2024)
Investigating Neural Audio Codecs for Speech Language Model-Based Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
von: Li, Jiaqi, et al.
Veröffentlicht: (2024)
Audio Deepfake Attribution: An Initial Dataset and Investigation
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
Language-Codec: Bridging Discrete Codec Representations and Speech Language Models
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
von: Ji, Shengpeng, et al.
Veröffentlicht: (2024)
SuperCodec: A Neural Speech Codec with Selective Back-Projection Network
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
von: Zheng, Youqiang, et al.
Veröffentlicht: (2024)
Neural Scoring: A Refreshed End-to-End Approach for Speaker Recognition in Complex Conditions
von: Lin, Wan, et al.
Veröffentlicht: (2024)
von: Lin, Wan, et al.
Veröffentlicht: (2024)
CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
von: Wang, Hankun, et al.
Veröffentlicht: (2025)
Probing the Robustness Properties of Neural Speech Codecs
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Interpreting End-to-End Deep Learning Models for Speech Source Localization Using Layer-wise Relevance Propagation
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
von: Comanducci, Luca, et al.
Veröffentlicht: (2024)
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
DualCodec: A Low-Frame-Rate, Semantically-Enhanced Neural Audio Codec for Speech Generation
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
von: Li, Jiaqi, et al.
Veröffentlicht: (2025)
Neural Codec-based Adversarial Sample Detection for Speaker Verification
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
von: Juvela, Lauri, et al.
Veröffentlicht: (2024)
PURE Codec: Progressive Unfolding of Residual Entropy for Speech Codec Learning
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
von: Shi, Jiatong, et al.
Veröffentlicht: (2025)
End-to-End Diarization utilizing Attractor Deep Clustering
von: Palzer, David, et al.
Veröffentlicht: (2025)
von: Palzer, David, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023) -
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024) -
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024) -
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023) -
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)