ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
Fuente:
arXiv
Salvato in:
| Autori principali: | Chou, Hsing-Hang, Lin, Yun-Shao, Sung, Ching-Chin, Tsao, Yu, Lee, Chi-Chun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-Shot Duet Singing Voices Separation with Diffusion Models
di: Yu, Chin-Yun, et al.
Pubblicazione: (2023)
di: Yu, Chin-Yun, et al.
Pubblicazione: (2023)
Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
di: Chen, Yun, et al.
Pubblicazione: (2023)
di: Chen, Yun, et al.
Pubblicazione: (2023)
REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers
di: Jiang, Yuepeng, et al.
Pubblicazione: (2025)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2025)
Towards Better Disentanglement in Non-Autoregressive Zero-Shot Expressive Voice Conversion
di: Akti, Seymanur, et al.
Pubblicazione: (2025)
di: Akti, Seymanur, et al.
Pubblicazione: (2025)
Discl-VC: Disentangled Discrete Tokens and In-Context Learning for Controllable Zero-Shot Voice Conversion
di: Wang, Kaidi, et al.
Pubblicazione: (2025)
di: Wang, Kaidi, et al.
Pubblicazione: (2025)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
di: Zhu, Xinfa, et al.
Pubblicazione: (2025)
CO-VADA: A Confidence-Oriented Voice Augmentation Debiasing Approach for Fair Speech Emotion Recognition
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
di: Tsai, Yun-Shao, et al.
Pubblicazione: (2025)
QR-VC: Leveraging Quantization Residuals for Linear Disentanglement in Zero-Shot Voice Conversion
di: Sim, Youngjun, et al.
Pubblicazione: (2024)
di: Sim, Youngjun, et al.
Pubblicazione: (2024)
Zero-Shot Sing Voice Conversion: built upon clustering-based phoneme representations
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
di: Zhou, Wangjin, et al.
Pubblicazione: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
Emo-bias: A Large Scale Evaluation of Social Bias on Speech Emotion Recognition
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2024)
EMO-Codec: An In-Depth Look at Emotion Preservation capacity of Legacy and Neural Codec Models With Subjective and Objective Evaluations
di: Ren, Wenze, et al.
Pubblicazione: (2024)
di: Ren, Wenze, et al.
Pubblicazione: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
di: Li, Yuke, et al.
Pubblicazione: (2024)
di: Li, Yuke, et al.
Pubblicazione: (2024)
GenVC: Self-Supervised Zero-Shot Voice Conversion
di: Cai, Zexin, et al.
Pubblicazione: (2025)
di: Cai, Zexin, et al.
Pubblicazione: (2025)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
di: Lee, Philip H., et al.
Pubblicazione: (2024)
di: Lee, Philip H., et al.
Pubblicazione: (2024)
Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
di: Zhang, Xueyao, et al.
Pubblicazione: (2025)
LDM-SVC: Latent Diffusion Model Based Zero-Shot Any-to-Any Singing Voice Conversion with Singer Guidance
di: Chen, Shihao, et al.
Pubblicazione: (2024)
di: Chen, Shihao, et al.
Pubblicazione: (2024)
Zero Shot Audio to Audio Emotion Transfer With Speaker Disentanglement
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
di: Dutta, Soumya, et al.
Pubblicazione: (2024)
Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
di: Zuo, Jialong, et al.
Pubblicazione: (2025)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
di: Yao, Jixun, et al.
Pubblicazione: (2024)
di: Yao, Jixun, et al.
Pubblicazione: (2024)
ControlVC: Zero-Shot Voice Conversion with Time-Varying Controls on Pitch and Speed
di: Chen, Meiying, et al.
Pubblicazione: (2022)
di: Chen, Meiying, et al.
Pubblicazione: (2022)
MeanVC: Lightweight and Streaming Zero-Shot Voice Conversion via Mean Flows
di: Ma, Guobin, et al.
Pubblicazione: (2025)
di: Ma, Guobin, et al.
Pubblicazione: (2025)
Everyone-Can-Sing: Zero-Shot Singing Voice Synthesis and Conversion with Speech Reference
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
di: Dai, Shuqi, et al.
Pubblicazione: (2025)
MaskVCT: Masked Voice Codec Transformer for Zero-Shot Voice Conversion With Increased Controllability via Multiple Guidances
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
di: Lee, Junhyeok, et al.
Pubblicazione: (2025)
Zero-shot Voice Conversion with Diffusion Transformers
di: Liu, Songting
Pubblicazione: (2024)
di: Liu, Songting
Pubblicazione: (2024)
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
Improvement Speaker Similarity for Zero-Shot Any-to-Any Voice Conversion of Whispered and Regular Speech
di: Avdeeva, Anastasia, et al.
Pubblicazione: (2024)
di: Avdeeva, Anastasia, et al.
Pubblicazione: (2024)
OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models
di: Zhu, Han, et al.
Pubblicazione: (2026)
di: Zhu, Han, et al.
Pubblicazione: (2026)
RT-VC: Real-Time Zero-Shot Voice Conversion with Speech Articulatory Coding
di: Liu, Yisi, et al.
Pubblicazione: (2025)
di: Liu, Yisi, et al.
Pubblicazione: (2025)
ASR for Affective Speech: Investigating Impact of Emotion and Speech Generative Strategy
di: Wu, Ya-Tse, et al.
Pubblicazione: (2026)
di: Wu, Ya-Tse, et al.
Pubblicazione: (2026)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
AdaLTM: Adaptive Layer-wise Task Vector Merging for Categorical Speech Emotion Recognition with ASR Knowledge Integration
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
di: Lee, Chia-Yu, et al.
Pubblicazione: (2026)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
Residual Speaker Representation for One-Shot Voice Conversion
di: Xu, Le, et al.
Pubblicazione: (2023)
di: Xu, Le, et al.
Pubblicazione: (2023)
Neural Concatenative Singing Voice Conversion: Rethinking Concatenation-Based Approach for One-Shot Singing Voice Conversion
di: Sha, Binzhu, et al.
Pubblicazione: (2023)
di: Sha, Binzhu, et al.
Pubblicazione: (2023)
Conan: A Chunkwise Online Network for Zero-Shot Adaptive Voice Conversion
di: Zhang, Yu, et al.
Pubblicazione: (2025)
di: Zhang, Yu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Zero-Shot Duet Singing Voices Separation with Diffusion Models
di: Yu, Chin-Yun, et al.
Pubblicazione: (2023) -
Lessons Learnt: Revisit Key Training Strategies for Effective Speech Emotion Recognition in the Wild
di: Tzeng, Jing-Tong, et al.
Pubblicazione: (2025) -
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2025) -
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
di: Chen, Zhengyang, et al.
Pubblicazione: (2024) -
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
di: Chen, Yun, et al.
Pubblicazione: (2023)