Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Qi, Tianhua, Wang, Shiyan, Lu, Cheng, Zhao, Yan, Zong, Yuan, Zheng, Wenming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
von: Qi, Tianhua, et al.
Veröffentlicht: (2025)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
von: Qi, Tianhua, et al.
Veröffentlicht: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
von: Kim, Sungjae, et al.
Veröffentlicht: (2022)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2024)
Improving Speaker-independent Speech Emotion Recognition Using Dynamic Joint Distribution Adaptation
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
von: Lu, Cheng, et al.
Veröffentlicht: (2024)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
von: Murata, Masato, et al.
Veröffentlicht: (2025)
von: Murata, Masato, et al.
Veröffentlicht: (2025)
EmotionCaps: Enhancing Audio Captioning Through Emotion-Augmented Data Generation
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
von: Manivannan, Mithun, et al.
Veröffentlicht: (2024)
Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
von: Chou, Huang-Cheng, et al.
Veröffentlicht: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
von: Choi, Ha-Yeong, et al.
Veröffentlicht: (2025)
Automatic Voice Classification Of Autistic Subjects
von: Vacca, Jessica, et al.
Veröffentlicht: (2024)
von: Vacca, Jessica, et al.
Veröffentlicht: (2024)
Generating Novel and Realistic Speakers for Voice Conversion
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
von: Chen, Meiying Melissa, et al.
Veröffentlicht: (2025)
Singing Voice Graph Modeling for SingFake Detection
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
von: Chen, Xuanjun, et al.
Veröffentlicht: (2024)
Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
von: Gao, Xiaoxue, et al.
Veröffentlicht: (2025)
Phoneme Discretized Saliency Maps for Explainable Detection of AI-Generated Voice
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
von: Gupta, Shubham, et al.
Veröffentlicht: (2024)
Neural Tracking of Sustained Attention, Attention Switching, and Natural Conversation in Audiovisual Environments using Mobile EEG
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
von: Wilroth, Johanna, et al.
Veröffentlicht: (2026)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
von: Ting, Zhu, et al.
Veröffentlicht: (2024)
Towards High-Quality and Efficient Speech Bandwidth Extension with Parallel Amplitude and Phase Prediction
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2024)
DeWinder: Single-Channel Wind Noise Reduction using Ultrasound Sensing
von: Yuan, Kuang, et al.
Veröffentlicht: (2024)
von: Yuan, Kuang, et al.
Veröffentlicht: (2024)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Toward Universal Speech Enhancement for Diverse Input Conditions
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
von: Zhang, Wangyou, et al.
Veröffentlicht: (2023)
Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
von: Yuan, Ze, et al.
Veröffentlicht: (2024)
Directional Selective Fixed-Filter Active Noise Control Based on a Convolutional Neural Network in Reverberant Environments
von: Wang, Boxiang, et al.
Veröffentlicht: (2026)
von: Wang, Boxiang, et al.
Veröffentlicht: (2026)
Align-ULCNet: Towards Low-Complexity and Robust Acoustic Echo and Noise Reduction
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
von: Shetu, Shrishti Saha, et al.
Veröffentlicht: (2024)
Soundscape Captioning using Sound Affective Quality Network and Large Language Model
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
von: Hou, Yuanbo, et al.
Veröffentlicht: (2024)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
von: Chen, Yun, et al.
Veröffentlicht: (2023)
von: Chen, Yun, et al.
Veröffentlicht: (2023)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Musical Score Following using Statistical Inference
von: Cowley, Josephine
Veröffentlicht: (2025)
von: Cowley, Josephine
Veröffentlicht: (2025)
Towards Improved Objective Perceptual Audio Quality Assessment -- Part 1: A Novel Data-Driven Cognitive Model
von: Delgado, Pablo M., et al.
Veröffentlicht: (2024)
von: Delgado, Pablo M., et al.
Veröffentlicht: (2024)
Dynamic Prediction of Full-Ocean Depth SSP by Hierarchical LSTM: An Experimental Result
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
von: Lu, Jiajun, et al.
Veröffentlicht: (2023)
Audio signal interpolation using optimal transportation of spectrograms
von: Valdivia, David, et al.
Veröffentlicht: (2025)
von: Valdivia, David, et al.
Veröffentlicht: (2025)
Self-supervised speech representation and contextual text embedding for match-mismatch classification with EEG recording
von: Wang, Bo, et al.
Veröffentlicht: (2024)
von: Wang, Bo, et al.
Veröffentlicht: (2024)
Adaptive Diagonal Loading using Krylov Subspaces for Robust Beamforming
von: Mittal, Manan, et al.
Veröffentlicht: (2026)
von: Mittal, Manan, et al.
Veröffentlicht: (2026)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
von: Wang, Kuan-Chen, et al.
Veröffentlicht: (2024)
Time-domain sound field estimation using kernel ridge regression
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
von: Brunnström, Jesper, et al.
Veröffentlicht: (2025)
STNet: Prediction of Underwater Sound Speed Profiles with An Advanced Semi-Transformer Neural Network
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
CochCeps-Augment: A Novel Self-Supervised Contrastive Learning Using Cochlear Cepstrum-based Masking for Speech Emotion Recognition
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
von: Ziogas, Ioannis, et al.
Veröffentlicht: (2024)
Computational Analysis of Yaredawi YeZema Silt in Ethiopian Orthodox Tewahedo Church Chants
von: Muluneh, Mequanent Argaw, et al.
Veröffentlicht: (2024)
von: Muluneh, Mequanent Argaw, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
von: Qi, Tianhua, et al.
Veröffentlicht: (2025) -
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
von: Qi, Tianhua, et al.
Veröffentlicht: (2024) -
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026) -
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024) -
EMOCONV-DIFF: Diffusion-based Speech Emotion Conversion for Non-parallel and In-the-wild Data
von: Prabhu, Navin Raj, et al.
Veröffentlicht: (2023)