USM-VC: Mitigating Timbre Leakage with Universal Semantic Mapping Residual Block for Voice Conversion
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Na, Wang, Chuke, Gu, Yu, Li, Zhifeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
von: Yang, Yuguang, et al.
Veröffentlicht: (2024)
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
von: Sim, Youngjun, et al.
Veröffentlicht: (2024)
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024)
von: Pan, Yu, et al.
Veröffentlicht: (2024)
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
von: Yao, Wenhan, et al.
Veröffentlicht: (2024)
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
von: Wang, Kaidi, et al.
Veröffentlicht: (2025)
EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
von: Joglekar, Advait, et al.
Veröffentlicht: (2025)
SelfVC: Voice Conversion With Iterative Refinement using Self Transformations
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
von: Neekhara, Paarth, et al.
Veröffentlicht: (2023)
The Voice Timbre Attribute Detection 2025 Challenge Evaluation Plan
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
von: Sheng, Zhengyan, et al.
Veröffentlicht: (2025)
Do Joint Language-Audio Embeddings Encode Perceptual Timbre Semantics?
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
von: Deng, Qixin, et al.
Veröffentlicht: (2025)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
SPA-SVC: Self-supervised Pitch Augmentation for Singing Voice Conversion
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
von: Bai, Bingsong, et al.
Veröffentlicht: (2024)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
von: Yutani, Tsugumasa, et al.
Veröffentlicht: (2024)
von: Yutani, Tsugumasa, et al.
Veröffentlicht: (2024)
Application of ASV for Voice Identification after VC and Duration Predictor Improvement in TTS Models
von: Nikolayevich, Borodin Kirill, et al.
Veröffentlicht: (2024)
von: Nikolayevich, Borodin Kirill, et al.
Veröffentlicht: (2024)
Diffusion Timbre Transfer Via Mutual Information Guided Inpainting
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
von: Lee, Ching Ho, et al.
Veröffentlicht: (2026)
Pureformer-VC: Non-parallel Voice Conversion with Pure Stylized Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
von: Yao, Wenhan, et al.
Veröffentlicht: (2025)
HQ-SVC: Towards High-Quality Zero-Shot Singing Voice Conversion in Low-Resource Scenarios
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
von: Bai, Bingsong, et al.
Veröffentlicht: (2025)
Prosody-Adaptable Audio Codecs for Zero-Shot Voice Conversion via In-Context Learning
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2025)
LHQ-SVC: Lightweight and High Quality Singing Voice Conversion Modeling
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
von: Huang, Yubo, et al.
Veröffentlicht: (2024)
PseudoVC: Improving One-shot Voice Conversion with Pseudo Paired Data
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
von: Cao, Songjun, et al.
Veröffentlicht: (2025)
Vec-Tok-VC+: Residual-enhanced Robust Zero-shot Voice Conversion with Progressive Constraints in a Dual-mode Training Strategy
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
von: Ma, Linhan, et al.
Veröffentlicht: (2024)
USM-Lite: Quantization and Sparsity Aware Fine-tuning for Speech Recognition with Universal Speech Models
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
von: Ding, Shaojin, et al.
Veröffentlicht: (2023)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
von: Li, Yuke, et al.
Veröffentlicht: (2024)
von: Li, Yuke, et al.
Veröffentlicht: (2024)
USM RNN-T model weights binarization
von: Rybakov, Oleg, et al.
Veröffentlicht: (2024)
von: Rybakov, Oleg, et al.
Veröffentlicht: (2024)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
von: Chen, Yun, et al.
Veröffentlicht: (2023)
von: Chen, Yun, et al.
Veröffentlicht: (2023)
Generative Adversarial Network based Voice Conversion: Techniques, Challenges, and Recent Advancements
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
von: Dhar, Sandipan, et al.
Veröffentlicht: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
von: Akti, Seymanur, et al.
Veröffentlicht: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
von: Du, Zhihao, et al.
Veröffentlicht: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
von: Li, Renyuan, et al.
Veröffentlicht: (2025)
Spectral Mapping of Singing Voices: U-Net-Assisted Vocal Segmentation
von: Sorrenti, Adam
Veröffentlicht: (2024)
von: Sorrenti, Adam
Veröffentlicht: (2024)
ReFlow-VC: Zero-shot Voice Conversion Based on Rectified Flow and Speaker Feature Optimization
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
von: Ren, Pengyu, et al.
Veröffentlicht: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
von: Byun, Kyungguen, et al.
Veröffentlicht: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
von: Saito, Yuki, et al.
Veröffentlicht: (2024)
Fast-VGAN: Lightweight Voice Conversion with Explicit Control of F0 and Duration Parameters
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
von: Abrassart, Mathilde, et al.
Veröffentlicht: (2025)
Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks
von: Chinchmalatpure, Prajwal, et al.
Veröffentlicht: (2025)
von: Chinchmalatpure, Prajwal, et al.
Veröffentlicht: (2025)
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
von: Gong, Yitian, et al.
Veröffentlicht: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
von: Yao, Jixun, et al.
Veröffentlicht: (2024)
VoiceShop: A Unified Speech-to-Speech Framework for Identity-Preserving Zero-Shot Voice Editing
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
von: Anastassiou, Philip, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Takin-VC: Expressive Zero-Shot Voice Conversion via Adaptive Hybrid Content Encoding and Enhanced Timbre Modeling
von: Yang, Yuguang, et al.
Veröffentlicht: (2024) -
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
von: Sim, Youngjun, et al.
Veröffentlicht: (2024) -
Zero-Shot Voice Conversion via Content-Aware Timbre Ensemble and Conditional Flow Matching
von: Pan, Yu, et al.
Veröffentlicht: (2024) -
Pureformer-VC: Non-parallel One-Shot Voice Conversion with Pure Transformer Blocks and Triplet Discriminative Training
von: Yao, Wenhan, et al.
Veröffentlicht: (2024) -
MulliVC: Multi-lingual Voice Conversion With Cycle Consistency
von: Huang, Jiawei, et al.
Veröffentlicht: (2024)