UDDETTS: Unifying Discrete and Dimensional Emotions for Controllable Emotional Text-to-Speech
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Jiaxuan, Xiang, Yang, Zhao, Han, Li, Xiangang, Gao, Yingying, Zhang, Shilei, Ling, Zhenhua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024)
FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes
di: Liu, Jiaxuan, et al.
Pubblicazione: (2026)
di: Liu, Jiaxuan, et al.
Pubblicazione: (2026)
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes
di: Li, Bocheng, et al.
Pubblicazione: (2025)
di: Li, Bocheng, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
di: R, Aron, et al.
Pubblicazione: (2024)
di: R, Aron, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with Distilled Prosodic and Linguistic Affect Representations
di: Shome, Debaditya, et al.
Pubblicazione: (2023)
di: Shome, Debaditya, et al.
Pubblicazione: (2023)
Emotion Entanglement and Bayesian Inference for Multi-Dimensional Emotion Understanding
di: Kotaprolu, Hemanth, et al.
Pubblicazione: (2026)
di: Kotaprolu, Hemanth, et al.
Pubblicazione: (2026)
EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
di: Yang, Guanrou, et al.
Pubblicazione: (2025)
DiffuSpeech: Silent Thought, Spoken Answer via Unified Speech-Text Diffusion
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
Beyond Global Emotion: Fine-Grained Emotional Speech Synthesis with Dynamic Word-Level Modulation
di: Wang, Sirui, et al.
Pubblicazione: (2025)
di: Wang, Sirui, et al.
Pubblicazione: (2025)
Emotion-Aligned Generation in Diffusion Text to Speech Models via Preference-Guided Optimization
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
di: Shi, Jiacheng, et al.
Pubblicazione: (2025)
Emotion is Not Just a Label: Latent Emotional Factors in LLM Processing
di: Reichman, Benjamin, et al.
Pubblicazione: (2026)
di: Reichman, Benjamin, et al.
Pubblicazione: (2026)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
di: Toyin, Hawau Olamide, et al.
Pubblicazione: (2024)
di: Toyin, Hawau Olamide, et al.
Pubblicazione: (2024)
Affective Multimodal Agents with Proactive Knowledge Grounding for Emotionally Aligned Marketing Dialogue
di: Yu, Lin, et al.
Pubblicazione: (2025)
di: Yu, Lin, et al.
Pubblicazione: (2025)
UniMoT: Unified Molecule-Text Language Model with Discrete Token Representation
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
di: Guo, Shuhan, et al.
Pubblicazione: (2024)
Speech Emotion Recognition Leveraging OpenAI's Whisper Representations and Attentive Pooling Methods
di: Shendabadi, Ali, et al.
Pubblicazione: (2026)
di: Shendabadi, Ali, et al.
Pubblicazione: (2026)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
Detecting Hope, Hate, and Emotion in Arabic Textual Speech and Multi-modal Memes Using Large Language Models
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
di: AlDahoul, Nouar, et al.
Pubblicazione: (2025)
Emergence of Hierarchical Emotion Organization in Large Language Models
di: Zhao, Bo, et al.
Pubblicazione: (2025)
di: Zhao, Bo, et al.
Pubblicazione: (2025)
A Review of Human Emotion Synthesis Based on Generative Technology
di: Ma, Fei, et al.
Pubblicazione: (2024)
di: Ma, Fei, et al.
Pubblicazione: (2024)
Emotion Detection From Social Media Posts
di: Rahman, Md Mahbubur, et al.
Pubblicazione: (2023)
di: Rahman, Md Mahbubur, et al.
Pubblicazione: (2023)
Emotion Classification in Low and Moderate Resource Languages
di: Tafreshi, Shabnam, et al.
Pubblicazione: (2024)
di: Tafreshi, Shabnam, et al.
Pubblicazione: (2024)
ReCode: Unify Plan and Action for Universal Granularity Control
di: Yu, Zhaoyang, et al.
Pubblicazione: (2025)
di: Yu, Zhaoyang, et al.
Pubblicazione: (2025)
Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
di: Liu, Rui, et al.
Pubblicazione: (2025)
di: Liu, Rui, et al.
Pubblicazione: (2025)
Large Language Models for Cross-lingual Emotion Detection
di: Kadiyala, Ram Mohan Rao
Pubblicazione: (2024)
di: Kadiyala, Ram Mohan Rao
Pubblicazione: (2024)
Generative Emotion Cause Explanation in Multimodal Conversations
di: Wang, Lin, et al.
Pubblicazione: (2024)
di: Wang, Lin, et al.
Pubblicazione: (2024)
A Simple Attention-Based Mechanism for Bimodal Emotion Classification
di: Elabd, Mazen, et al.
Pubblicazione: (2024)
di: Elabd, Mazen, et al.
Pubblicazione: (2024)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
di: Tsujimura, Hikaru, et al.
Pubblicazione: (2025)
AcademicEval: Live Long-Context LLM Benchmark
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
di: Zhang, Haozhen, et al.
Pubblicazione: (2025)
Improving Direct Persian-English Speech-to-Speech Translation with Discrete Units and Synthetic Parallel Data
di: Rashidi, Sina, et al.
Pubblicazione: (2025)
di: Rashidi, Sina, et al.
Pubblicazione: (2025)
Controlled Generation for Private Synthetic Text
di: Zhao, Zihao, et al.
Pubblicazione: (2025)
di: Zhao, Zihao, et al.
Pubblicazione: (2025)
Semantic Differentiation in Speech Emotion Recognition: Insights from Descriptive and Expressive Speech Roles
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
di: Guo, Rongchen, et al.
Pubblicazione: (2025)
ELSA: A Style Aligned Dataset for Emotionally Intelligent Language Generation
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
di: Gandhi, Vishal, et al.
Pubblicazione: (2025)
Evaluating the Effectiveness of Data Augmentation for Emotion Classification in Low-Resource Settings
di: Arora, Aashish, et al.
Pubblicazione: (2024)
di: Arora, Aashish, et al.
Pubblicazione: (2024)
TED: Turn Emphasis with Dialogue Feature Attention for Emotion Recognition in Conversation
di: Ono, Junya, et al.
Pubblicazione: (2025)
di: Ono, Junya, et al.
Pubblicazione: (2025)
EmoLLM: Appraisal-Grounded Cognitive-Emotional Co-Reasoning in Large Language Models
di: Zhang, Yifei, et al.
Pubblicazione: (2026)
di: Zhang, Yifei, et al.
Pubblicazione: (2026)
Language-Specific Representation of Emotion-Concept Knowledge Causally Supports Emotion Inference
di: Li, Ming, et al.
Pubblicazione: (2023)
di: Li, Ming, et al.
Pubblicazione: (2023)
MoST: Mixing Speech and Text with Modality-Aware Mixture of Experts
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
di: Lou, Yuxuan, et al.
Pubblicazione: (2026)
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
di: Attaluri, Kaushal, et al.
Pubblicazione: (2024)
di: Attaluri, Kaushal, et al.
Pubblicazione: (2024)
SER Evals: In-domain and Out-of-domain Benchmarking for Speech Emotion Recognition
di: Osman, Mohamed, et al.
Pubblicazione: (2024)
di: Osman, Mohamed, et al.
Pubblicazione: (2024)
Documenti analoghi
-
DiffStyleTTS: Diffusion-based Hierarchical Prosody Modeling for Text-to-Speech with Diverse and Controllable Styles
di: Liu, Jiaxuan, et al.
Pubblicazione: (2024) -
FunCineForge: A Unified Dataset Toolkit and Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes
di: Liu, Jiaxuan, et al.
Pubblicazione: (2026) -
Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion Processes
di: Li, Bocheng, et al.
Pubblicazione: (2025) -
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023) -
SEGAA: A Unified Approach to Predicting Age, Gender, and Emotion in Speech
di: R, Aron, et al.
Pubblicazione: (2024)