Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Yang, Haoyuan, Yang, Mu, Xie, Jiamin, Chen, Szu-Jui, Hansen, John H. L. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
por: Yang, Mu, et al.
Publicado: (2025)
por: Yang, Mu, et al.
Publicado: (2025)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026)
por: Chen, Szu-Jui, et al.
Publicado: (2026)
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
por: Qi, Tianhua, et al.
Publicado: (2024)
por: Qi, Tianhua, et al.
Publicado: (2024)
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
por: Qi, Tianhua, et al.
Publicado: (2025)
por: Qi, Tianhua, et al.
Publicado: (2025)
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
por: Chen, Yun, et al.
Publicado: (2023)
por: Chen, Yun, et al.
Publicado: (2023)
Maestro-EVC: Controllable Emotional Voice Conversion Guided by References and Explicit Prosody
por: Yoon, Jinsung, et al.
Publicado: (2025)
por: Yoon, Jinsung, et al.
Publicado: (2025)
ZSDEVC: Zero-Shot Diffusion-based Emotional Voice Conversion with Disentangled Mechanism
por: Chou, Hsing-Hang, et al.
Publicado: (2024)
por: Chou, Hsing-Hang, et al.
Publicado: (2024)
Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
por: Yang, Mu, et al.
Publicado: (2026)
por: Yang, Mu, et al.
Publicado: (2026)
MixRep: Hidden Representation Mixup for Low-Resource Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2023)
por: Xie, Jiamin, et al.
Publicado: (2023)
DEFORMER: Coupling Deformed Localized Patterns with Global Context for Robust End-to-end Speech Recognition
por: Xie, Jiamin, et al.
Publicado: (2022)
por: Xie, Jiamin, et al.
Publicado: (2022)
MuSE-SVS: Multi-Singer Emotional Singing Voice Synthesizer that Controls Emotional Intensity
por: Kim, Sungjae, et al.
Publicado: (2022)
por: Kim, Sungjae, et al.
Publicado: (2022)
StreamVoice: Streamable Context-Aware Language Modeling for Real-time Zero-Shot Voice Conversion
por: Wang, Zhichao, et al.
Publicado: (2024)
por: Wang, Zhichao, et al.
Publicado: (2024)
PAVITS: Exploring Prosody-aware VITS for End-to-End Emotional Voice Conversion
por: Qi, Tianhua, et al.
Publicado: (2024)
por: Qi, Tianhua, et al.
Publicado: (2024)
E-chat: Emotion-sensitive Spoken Dialogue System with Large Language Models
por: Xue, Hongfei, et al.
Publicado: (2023)
por: Xue, Hongfei, et al.
Publicado: (2023)
Enhancing In-the-Wild Speech Emotion Conversion with Resynthesis-based Duration Modeling
por: Prabhu, Navin Raj, et al.
Publicado: (2025)
por: Prabhu, Navin Raj, et al.
Publicado: (2025)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
por: Du, Zongyang, et al.
Publicado: (2025)
por: Du, Zongyang, et al.
Publicado: (2025)
S$^2$Voice: Style-Aware Autoregressive Modeling with Enhanced Conditioning for Singing Style Conversion
por: Wang, Ziqian, et al.
Publicado: (2026)
por: Wang, Ziqian, et al.
Publicado: (2026)
ClapFM-EVC: High-Fidelity and Flexible Emotional Voice Conversion with Dual Control from Natural Language and Speech
por: Pan, Yu, et al.
Publicado: (2025)
por: Pan, Yu, et al.
Publicado: (2025)
Towards Naturalistic Voice Conversion: NaturalVoices Dataset with an Automatic Processing Pipeline
por: Salman, Ali N., et al.
Publicado: (2024)
por: Salman, Ali N., et al.
Publicado: (2024)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
por: Yang, Guanrou, et al.
Publicado: (2025)
por: Yang, Guanrou, et al.
Publicado: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
por: Yao, Wenhan, et al.
Publicado: (2024)
por: Yao, Wenhan, et al.
Publicado: (2024)
StableVC: Style Controllable Zero-Shot Voice Conversion with Conditional Flow Matching
por: Yao, Jixun, et al.
Publicado: (2024)
por: Yao, Jixun, et al.
Publicado: (2024)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
por: Tu, Wenming, et al.
Publicado: (2025)
por: Tu, Wenming, et al.
Publicado: (2025)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
por: Yang, Yiqing, et al.
Publicado: (2025)
por: Yang, Yiqing, et al.
Publicado: (2025)
EmoOmni: Bridging Emotional Understanding and Expression in Omni-Modal LLMs
por: Tian, Wenjie, et al.
Publicado: (2026)
por: Tian, Wenjie, et al.
Publicado: (2026)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
por: Wang, Zhichao, et al.
Publicado: (2024)
por: Wang, Zhichao, et al.
Publicado: (2024)
Speaker-agnostic Emotion Vector for Cross-speaker Emotion Intensity Control
por: Murata, Masato, et al.
Publicado: (2025)
por: Murata, Masato, et al.
Publicado: (2025)
Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
por: Zhou, Wangzixi, et al.
Publicado: (2026)
por: Zhou, Wangzixi, et al.
Publicado: (2026)
DreamVoice: Text-Guided Voice Conversion
por: Hai, Jiarui, et al.
Publicado: (2024)
por: Hai, Jiarui, et al.
Publicado: (2024)
EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing
por: Cong, Gaoxiang, et al.
Publicado: (2024)
por: Cong, Gaoxiang, et al.
Publicado: (2024)
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation
por: Zuo, Tianlun, et al.
Publicado: (2025)
por: Zuo, Tianlun, et al.
Publicado: (2025)
OneVoice: One Model, Triple Scenarios-Towards Unified Zero-shot Voice Conversion
por: Wang, Zhichao, et al.
Publicado: (2026)
por: Wang, Zhichao, et al.
Publicado: (2026)
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions
por: Chou, Huang-Cheng, et al.
Publicado: (2025)
por: Chou, Huang-Cheng, et al.
Publicado: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
por: Wang, Tianrui, et al.
Publicado: (2025)
por: Wang, Tianrui, et al.
Publicado: (2025)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
por: Zhao, Yan, et al.
Publicado: (2024)
por: Zhao, Yan, et al.
Publicado: (2024)
Emotion-Aware Speech Generation with Character-Specific Voices for Comics
por: Qian, Zhiwen, et al.
Publicado: (2025)
por: Qian, Zhiwen, et al.
Publicado: (2025)
Enhancing Spectrogram Realism in Singing Voice Synthesis via Explicit Bandwidth Extension Prior to Vocoder
por: Yang, Runxuan, et al.
Publicado: (2025)
por: Yang, Runxuan, et al.
Publicado: (2025)
Acoustic modeling for Overlapping Speech Recognition: JHU Chime-5 Challenge System
por: Manohar, Vimal, et al.
Publicado: (2024)
por: Manohar, Vimal, et al.
Publicado: (2024)
LLM supervised Pre-training for Multimodal Emotion Recognition in Conversations
por: Dutta, Soumya, et al.
Publicado: (2025)
por: Dutta, Soumya, et al.
Publicado: (2025)
Emotion-Anchored Contrastive Learning Framework for Emotion Recognition in Conversation
por: Yu, Fangxu, et al.
Publicado: (2024)
por: Yu, Fangxu, et al.
Publicado: (2024)
Ejemplares similares
-
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
por: Yang, Mu, et al.
Publicado: (2025) -
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026) -
Towards Realistic Emotional Voice Conversion using Controllable Emotional Intensity
por: Qi, Tianhua, et al.
Publicado: (2024) -
PromptEVC: Controllable Emotional Voice Conversion with Natural Language Prompts
por: Qi, Tianhua, et al.
Publicado: (2025) -
Attention-based Interactive Disentangling Network for Instance-level Emotional Voice Conversion
por: Chen, Yun, et al.
Publicado: (2023)