Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Jing, Xin, Zhou, Kun, Triantafyllopoulos, Andreas, Schuller, Björn W. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
di: Jing, Xin, et al.
Pubblicazione: (2024)
di: Jing, Xin, et al.
Pubblicazione: (2024)
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
di: Jing, Xin, et al.
Pubblicazione: (2026)
di: Jing, Xin, et al.
Pubblicazione: (2026)
Abusive Speech Detection in Indic Languages Using Acoustic Features
di: Spiesberger, Anika A., et al.
Pubblicazione: (2024)
di: Spiesberger, Anika A., et al.
Pubblicazione: (2024)
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025)
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
di: Jing, Xin, et al.
Pubblicazione: (2025)
di: Jing, Xin, et al.
Pubblicazione: (2025)
Audio-based Step-count Estimation for Running -- Windowing and Neural Network Baselines
di: Wagner, Philipp, et al.
Pubblicazione: (2024)
di: Wagner, Philipp, et al.
Pubblicazione: (2024)
An automatic analysis of ultrasound vocalisations for the prediction of interaction context in captive Egyptian fruit bats
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
di: Zhao, Yan, et al.
Pubblicazione: (2024)
di: Zhao, Yan, et al.
Pubblicazione: (2024)
Computer Audition: From Task-Specific Machine Learning to Foundation Models
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2024)
Can Large Language Models Aid in Annotating Speech Emotional Data? Uncovering New Frontiers
di: Latif, Siddique, et al.
Pubblicazione: (2023)
di: Latif, Siddique, et al.
Pubblicazione: (2023)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
Testing Correctness, Fairness, and Robustness of Speech Emotion Recognition Models
di: Derington, Anna, et al.
Pubblicazione: (2023)
di: Derington, Anna, et al.
Pubblicazione: (2023)
Quantifying Dimensional Independence in Speech: An Information-Theoretic Framework for Disentangled Representation Learning
di: Kashyap, Bipasha, et al.
Pubblicazione: (2026)
di: Kashyap, Bipasha, et al.
Pubblicazione: (2026)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
autrainer: A Modular and Extensible Deep Learning Toolkit for Computer Audition Tasks
di: Rampp, Simon, et al.
Pubblicazione: (2024)
di: Rampp, Simon, et al.
Pubblicazione: (2024)
Domain Adapting Deep Reinforcement Learning for Real-world Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2022)
Audio Enhancement for Computer Audition -- An Iterative Training Paradigm Using Sample Importance
di: Milling, Manuel, et al.
Pubblicazione: (2024)
di: Milling, Manuel, et al.
Pubblicazione: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2023)
Representation Learning with Parameterised Quantum Circuits for Advancing Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2025)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2025)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024)
di: Bott, Thomas, et al.
Pubblicazione: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Exploring Meta Information for Audio-based Zero-shot Bird Classification
di: Gebhard, Alexander, et al.
Pubblicazione: (2023)
di: Gebhard, Alexander, et al.
Pubblicazione: (2023)
Breaking Resource Barriers in Speech Emotion Recognition via Data Distillation
di: Chang, Yi, et al.
Pubblicazione: (2024)
di: Chang, Yi, et al.
Pubblicazione: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
Intelligent Cardiac Auscultation for Murmur Detection via Parallel-Attentive Models with Uncertainty Estimation
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
Bringing the Discussion of Minima Sharpness to the Audio Domain: a Filter-Normalised Evaluation for Acoustic Scene Classification
di: Milling, Manuel, et al.
Pubblicazione: (2023)
di: Milling, Manuel, et al.
Pubblicazione: (2023)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
di: He, Xiangheng, et al.
Pubblicazione: (2024)
di: He, Xiangheng, et al.
Pubblicazione: (2024)
DOTA-ME-CS: Daily Oriented Text Audio-Mandarin English-Code Switching Dataset
di: Li, Yupei, et al.
Pubblicazione: (2025)
di: Li, Yupei, et al.
Pubblicazione: (2025)
Leveraging Local and Global Knowledge Integration with Time-Frequency Calibrated Distillation for Speech Enhancement
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
di: Cheng, Jiaming, et al.
Pubblicazione: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
di: Sankey-Olsen, Cuno, et al.
Pubblicazione: (2025)
Crab: Multi Layer Contrastive Supervision to Improve Speech Emotion Recognition Under Both Acted and Natural Speech Condition
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
di: Ueda, Lucas H., et al.
Pubblicazione: (2026)
Are you sure? Analysing Uncertainty Quantification Approaches for Real-world Speech Emotion Recognition
di: Schrüfer, Oliver, et al.
Pubblicazione: (2024)
di: Schrüfer, Oliver, et al.
Pubblicazione: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
di: Sun, Haoqin, et al.
Pubblicazione: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
di: Rajapakshe, Thejan, et al.
Pubblicazione: (2024)
Emotion-Coherent Speech Data Augmentation and Self-Supervised Contrastive Style Training for Enhancing Kids's Story Speech Synthesis
di: Chung, Raymond
Pubblicazione: (2026)
di: Chung, Raymond
Pubblicazione: (2026)
DualSpeech: Enhancing Speaker-Fidelity and Text-Intelligibility Through Dual Classifier-Free Guidance
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
di: Yang, Jinhyeok, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ParaCLAP -- Towards a general language-audio model for computational paralinguistic tasks
di: Jing, Xin, et al.
Pubblicazione: (2024) -
SmoothCLAP: Soft-Target Enhanced Contrastive Language\--Audio Pretraining for Affective Computing
di: Jing, Xin, et al.
Pubblicazione: (2026) -
Abusive Speech Detection in Indic Languages Using Acoustic Features
di: Spiesberger, Anika A., et al.
Pubblicazione: (2024) -
Charting 15 years of progress in deep learning for speech emotion recognition: A replication study
di: Triantafyllopoulos, Andreas, et al.
Pubblicazione: (2025) -
MELT: Towards Automated Multimodal Emotion Data Annotation by Leveraging LLM Embedded Knowledge
di: Jing, Xin, et al.
Pubblicazione: (2025)