Residual Speaker Representation for One-Shot Voice Conversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Le, Yi, Jiangyan, Wang, Tao, Ren, Yong, Zhong, Rongxiu, Wen, Zhengqi, Tao, Jianhua |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026)
von: Ren, Yong, et al.
Veröffentlicht: (2026)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023)
von: Ren, Yong, et al.
Veröffentlicht: (2023)
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)
Dual-Branch Knowledge Distillation for Noise-Robust Synthetic Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
Towards Robust Audio Deepfake Detection: A Evolving Benchmark for Continual Learning
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
von: Zhang, Xiaohui, et al.
Veröffentlicht: (2024)
WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity Verification
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024)
ALLM4ADD: Unlocking the Capabilities of Audio Large Language Models for Audio Deepfake Detection
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
von: Wang, Zhichao, et al.
Veröffentlicht: (2026)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
von: Sha, Binzhu, et al.
Veröffentlicht: (2023)
Utilizing Speaker Profiles for Impersonation Audio Detection
von: Gu, Hao, et al.
Veröffentlicht: (2024)
von: Gu, Hao, et al.
Veröffentlicht: (2024)
Distinguishing Neural Speech Synthesis Models Through Fingerprints in Speech Waveforms
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
von: Zhang, Chu Yuan, et al.
Veröffentlicht: (2023)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
ADD 2023: Towards Audio Deepfake Detection and Analysis in the Wild
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
von: Yi, Jiangyan, et al.
Veröffentlicht: (2024)
EmoFake: An Initial Dataset for Emotion Fake Audio Detection
von: Zhao, Yan, et al.
Veröffentlicht: (2022)
von: Zhao, Yan, et al.
Veröffentlicht: (2022)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
von: Avdeeva, Anastasia, et al.
Veröffentlicht: (2024)
Noro: Noise-Robust One-shot Voice Conversion with Hidden Speaker Representation Learning
von: He, Haorui, et al.
Veröffentlicht: (2024)
von: He, Haorui, et al.
Veröffentlicht: (2024)
JoyVoice: Long-Context Conditioning for Anthropomorphic Multi-Speaker Conversational Synthesis
von: Yu, Fan, et al.
Veröffentlicht: (2025)
von: Yu, Fan, et al.
Veröffentlicht: (2025)
Quantifying Source Speaker Leakage in One-to-One Voice Conversion
von: Wellington, Scott, et al.
Veröffentlicht: (2025)
von: Wellington, Scott, et al.
Veröffentlicht: (2025)
Audio Deepfake Attribution: An Initial Dataset and Investigation
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
von: Yan, Xinrui, et al.
Veröffentlicht: (2022)
Spatial Reconstructed Local Attention Res2Net with F0 Subband for Fake Speech Detection
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
von: Fan, Cunhang, et al.
Veröffentlicht: (2023)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
von: Wang, Zhichao, et al.
Veröffentlicht: (2023)
Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
von: Jeon, Yejin, et al.
Veröffentlicht: (2024)
An Unsupervised Domain Adaptation Method for Locating Manipulated Region in partially fake Audio
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
von: Zeng, Siding, et al.
Veröffentlicht: (2024)
Reject Threshold Adaptation for Open-Set Model Attribution of Deepfake Audio
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
von: Yan, Xinrui, et al.
Veröffentlicht: (2024)
REWIND: Speech Time Reversal for Enhancing Speaker Representations in Diffusion-based Voice Conversion
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
von: Biyani, Ishan D., et al.
Veröffentlicht: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
von: Wang, Zhichao, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
von: Zhou, Wangjin, et al.
Veröffentlicht: (2024)
SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention
von: Li, Junjie, et al.
Veröffentlicht: (2023)
von: Li, Junjie, et al.
Veröffentlicht: (2023)
Learning Emotion-Invariant Speaker Representations for Speaker Verification
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
von: Tian, Jingguang, et al.
Veröffentlicht: (2025)
PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
von: Shi, Shuchen, et al.
Veröffentlicht: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
von: Dai, Shuqi, et al.
Veröffentlicht: (2025)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
von: Ruggiero, Giuseppe, et al.
Veröffentlicht: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
von: Park, Nohil, et al.
Veröffentlicht: (2024)
von: Park, Nohil, et al.
Veröffentlicht: (2024)
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
RawBMamba: End-to-End Bidirectional State Space Model for Audio Deepfake Detection
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
von: Chen, Yujie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Edit Content, Preserve Acoustics: Imperceptible Text-Based Speech Editing via Self-Consistency Rewards
von: Ren, Yong, et al.
Veröffentlicht: (2026) -
OV-InstructTTS: Towards Open-Vocabulary Instruct Text-to-Speech
von: Ren, Yong, et al.
Veröffentlicht: (2026) -
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
von: Zhou, Junzuo, et al.
Veröffentlicht: (2024) -
Fewer-token Neural Speech Codec with Time-invariant Codes
von: Ren, Yong, et al.
Veröffentlicht: (2023) -
ASRRL-TTS: Agile Speaker Representation Reinforcement Learning for Text-to-Speech Speaker Adaptation
von: Fu, Ruibo, et al.
Veröffentlicht: (2024)