PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
Fuente:
arXiv
Salvato in:
| Autori principali: | Qi, Tianhua, Zheng, Wenming, Lu, Cheng, Zong, Yuan, Lian, Hailun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
di: Qi, Tianhua, et al.
Pubblicazione: (2024)
di: Qi, Tianhua, et al.
Pubblicazione: (2024)
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
di: Qi, Tianhua, et al.
Pubblicazione: (2025)
di: Qi, Tianhua, et al.
Pubblicazione: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
di: Lu, Cheng, et al.
Pubblicazione: (2024)
di: Lu, Cheng, et al.
Pubblicazione: (2024)
On Improving Error Resilience of Neural End-to-End Speech Coders
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
di: Gupta, Kishan, et al.
Pubblicazione: (2024)
AADNet: An End-to-End Deep Learning Model for Auditory Attention Decoding
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
di: Du, Zongyang, et al.
Pubblicazione: (2024)
di: Du, Zongyang, et al.
Pubblicazione: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
di: Kang, Wonjune, et al.
Pubblicazione: (2022)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
di: Zhao, Yan, et al.
Pubblicazione: (2024)
di: Zhao, Yan, et al.
Pubblicazione: (2024)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
di: Prabhu, Navin Raj, et al.
Pubblicazione: (2023)
SynthVC: Leveraging Synthetic Data for End-to-End Low Latency Streaming Voice Conversion
di: Guo, Zhao, et al.
Pubblicazione: (2025)
di: Guo, Zhao, et al.
Pubblicazione: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
di: Wang, Zhichao, et al.
Pubblicazione: (2024)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
di: Choi, Ha-Yeong, et al.
Pubblicazione: (2025)
di: Choi, Ha-Yeong, et al.
Pubblicazione: (2025)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
FunnelNet: An End-to-End Deep Learning Framework to Monitor Digital Heart Murmur in Real-Time
di: Jobayer, Md, et al.
Pubblicazione: (2024)
di: Jobayer, Md, et al.
Pubblicazione: (2024)
Automatic Voice Classification Of Autistic Subjects
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
di: Vacca, Jessica, et al.
Pubblicazione: (2024)
Singing Voice Graph Modeling for SingFake Detection
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
di: Chen, Xuanjun, et al.
Pubblicazione: (2024)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
di: Gupta, Shubham, et al.
Pubblicazione: (2024)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
Accent-VITS:accent transfer for end-to-end TTS
di: Ma, Linhan, et al.
Pubblicazione: (2023)
di: Ma, Linhan, et al.
Pubblicazione: (2023)
VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation
di: Yu, Yifeng, et al.
Pubblicazione: (2024)
di: Yu, Yifeng, et al.
Pubblicazione: (2024)
Disentangling the Prosody and Semantic Information with Pre-trained Model for In-Context Learning based Zero-Shot Voice Conversion
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
di: Wilroth, Johanna, et al.
Pubblicazione: (2026)
di: Wilroth, Johanna, et al.
Pubblicazione: (2026)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
di: Chou, Huang-Cheng, et al.
Pubblicazione: (2024)
Period Singer: Integrating Periodic and Aperiodic Variational Autoencoders for Natural-Sounding End-to-End Singing Voice Synthesis
di: Kim, Taewoo, et al.
Pubblicazione: (2024)
di: Kim, Taewoo, et al.
Pubblicazione: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
di: Li, Shaojun, et al.
Pubblicazione: (2024)
di: Li, Shaojun, et al.
Pubblicazione: (2024)
Dynamic Prediction of Full-Ocean Depth SSP by Hierarchical LSTM: An Experimental Result
di: Lu, Jiajun, et al.
Pubblicazione: (2023)
di: Lu, Jiajun, et al.
Pubblicazione: (2023)
Speaker Adaptation for Quantised End-to-End ASR Models
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
di: Yuan, Kuang, et al.
Pubblicazione: (2024)
di: Yuan, Kuang, et al.
Pubblicazione: (2024)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
di: Yuan, Ze, et al.
Pubblicazione: (2024)
di: Yuan, Ze, et al.
Pubblicazione: (2024)
End-to-End Integration of Speech Separation and Voice Activity Detection for Low-Latency Diarization of Telephone Conversations
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
di: Morrone, Giovanni, et al.
Pubblicazione: (2023)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
Future Full-Ocean Deep SSPs Prediction based on Hierarchical Long Short-Term Memory Neural Networks
di: Lu, Jiajun, et al.
Pubblicazione: (2023)
di: Lu, Jiajun, et al.
Pubblicazione: (2023)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
Frequency-Based Alignment of EEG and Audio Signals Using Contrastive Learning and SincNet for Auditory Attention Detection
di: Liao, Yuan, et al.
Pubblicazione: (2025)
di: Liao, Yuan, et al.
Pubblicazione: (2025)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
di: Yoon, Jinsung, et al.
Pubblicazione: (2025)
di: Yoon, Jinsung, et al.
Pubblicazione: (2025)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
di: Raj, Desh
Pubblicazione: (2024)
di: Raj, Desh
Pubblicazione: (2024)
SAML: Speaker Adaptive Mixture of LoRA Experts for End-to-End ASR
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
di: Qi, Tianhua, et al.
Pubblicazione: (2024) -
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
di: Qi, Tianhua, et al.
Pubblicazione: (2025) -
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026) -
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
di: Lu, Cheng, et al.
Pubblicazione: (2024) -
On Improving Error Resilience of Neural End-to-End Speech Coders
di: Gupta, Kishan, et al.
Pubblicazione: (2024)