Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Kun, Zhang, You, Ng, Dianwen, Zhao, Shengkui, Wang, Hao, Ma, Bin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023)
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
Dataset-Distillation Generative Model for Speech Emotion Recognition
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
von: Ritter-Gutierrez, Fabian, et al.
Veröffentlicht: (2024)
ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
HiFi-SR: A Unified Generative Transformer-Convolutional Adversarial Network for High-Fidelity Speech Super-Resolution
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)
SPGM: Prioritizing Local Features for enhanced speech separation performance
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
von: Yip, Jia Qi, et al.
Veröffentlicht: (2023)
FRCRN: Boosting Feature Representation using Frequency Recurrence for Monaural Speech Enhancement
von: Zhao, Shengkui, et al.
Veröffentlicht: (2022)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2022)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
Multi-band Frequency Reconstruction for Neural Psychoacoustic Coding
von: Ng, Dianwen, et al.
Veröffentlicht: (2025)
von: Ng, Dianwen, et al.
Veröffentlicht: (2025)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
von: Bott, Thomas, et al.
Veröffentlicht: (2024)
EMO-RL: Emotion-Rule-Based Reinforcement Learning Enhanced Audio-Language Model for Generalized Speech Emotion Recognition
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
von: Li, Pengcheng, et al.
Veröffentlicht: (2025)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
EmoSphere-TTS: Emotional Style and Intensity Modeling via Spherical Emotion Vector for Controllable Emotional Text-to-Speech
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
von: Cho, Deok-Hyeon, et al.
Veröffentlicht: (2024)
Online Audio-Visual Autoregressive Speaker Extraction
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
von: Pan, Zexu, et al.
Veröffentlicht: (2025)
Monaural Speech Enhancement with Complex Convolutional Block Attention Module and Joint Time Frequency Losses
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
von: Zhao, Shengkui, et al.
Veröffentlicht: (2021)
Adaptive Speech Emotion Representation Learning Based On Dynamic Graph
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
von: Gao, Yingxue, et al.
Veröffentlicht: (2024)
Temporal-Frequency State Space Duality: An Efficient Paradigm for Speech Emotion Recognition
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
von: Zhao, Jiaqi, et al.
Veröffentlicht: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Iterative Prototype Refinement for Ambiguous Speech Emotion Recognition
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
von: Sun, Haoqin, et al.
Veröffentlicht: (2024)
Emotional Text-To-Speech Based on Mutual-Information-Guided Emotion-Timbre Disentanglement
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
von: Yang, Jianing, et al.
Veröffentlicht: (2025)
Semantic-Emotional Resonance Embedding: A Semi-Supervised Paradigm for Cross-Lingual Speech Emotion Recognition
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
von: Zhao, Ya, et al.
Veröffentlicht: (2026)
Emotion-Aware Contrastive Adaptation Network for Source-Free Cross-Corpus Speech Emotion Recognition
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
von: Zhao, Yan, et al.
Veröffentlicht: (2024)
EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Speech Emotion Recognition with ASR Integration
von: Li, Yuanchao
Veröffentlicht: (2026)
von: Li, Yuanchao
Veröffentlicht: (2026)
AmbER$^2$: Dual Ambiguity-Aware Emotion Recognition Applied to Speech and Text
von: Wu, Jingyao, et al.
Veröffentlicht: (2026)
von: Wu, Jingyao, et al.
Veröffentlicht: (2026)
VECL-TTS: Voice identity and Emotional style controllable Cross-Lingual Text-to-Speech
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
von: Gudmalwar, Ashishkumar, et al.
Veröffentlicht: (2024)
RA-CLAP: Relation-Augmented Emotional Speaking Style Contrastive Language-Audio Pretraining For Speech Retrieval
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
von: Sun, Haoqin, et al.
Veröffentlicht: (2025)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
von: Dutta, Soumya, et al.
Veröffentlicht: (2025)
Machine Unlearning in Speech Emotion Recognition via Forget Set Alone
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
von: Ren, Zhao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
von: Zhou, Kun, et al.
Veröffentlicht: (2024) -
MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation
von: Zhao, Shengkui, et al.
Veröffentlicht: (2023) -
Towards Audio Codec-based Speech Separation
von: Yip, Jia Qi, et al.
Veröffentlicht: (2024) -
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
Conditional Latent Diffusion-Based Speech Enhancement Via Dual Context Learning
von: Zhao, Shengkui, et al.
Veröffentlicht: (2025)