PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Shaozuo, Mehrish, Ambuj, Li, Yingting, Poria, Soujanya |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
di: Li, Yingting, et al.
Pubblicazione: (2024)
di: Li, Yingting, et al.
Pubblicazione: (2024)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
di: Li, Yingting, et al.
Pubblicazione: (2024)
di: Li, Yingting, et al.
Pubblicazione: (2024)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Improving Text-To-Audio Models with Synthetic Captions
di: Kong, Zhifeng, et al.
Pubblicazione: (2024)
di: Kong, Zhifeng, et al.
Pubblicazione: (2024)
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
di: Hung, Chia-Yu, et al.
Pubblicazione: (2024)
di: Hung, Chia-Yu, et al.
Pubblicazione: (2024)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024)
di: Bott, Thomas, et al.
Pubblicazione: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
di: Peng, Yifan, et al.
Pubblicazione: (2024)
di: Peng, Yifan, et al.
Pubblicazione: (2024)
Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization
di: Majumder, Navonil, et al.
Pubblicazione: (2024)
di: Majumder, Navonil, et al.
Pubblicazione: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Phonetic Enhanced Language Modeling for Text-to-Speech Synthesis
di: Zhou, Kun, et al.
Pubblicazione: (2024)
di: Zhou, Kun, et al.
Pubblicazione: (2024)
Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback
di: Chen, Chen, et al.
Pubblicazione: (2024)
di: Chen, Chen, et al.
Pubblicazione: (2024)
Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
JELLY: Joint Emotion Recognition and Context Reasoning with LLMs for Conversational Speech Synthesis
di: Cha, Jun-Hyeok, et al.
Pubblicazione: (2025)
di: Cha, Jun-Hyeok, et al.
Pubblicazione: (2025)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
ProsodyFM: Unsupervised Phrasing and Intonation Control for Intelligible Speech Synthesis
di: He, Xiangheng, et al.
Pubblicazione: (2024)
di: He, Xiangheng, et al.
Pubblicazione: (2024)
Expressive Prompting: Improving Emotion Intensity and Speaker Consistency in Zero-Shot TTS
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
di: Kato, Shuhei
Pubblicazione: (2025)
di: Kato, Shuhei
Pubblicazione: (2025)
Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework
di: Li, Zhu, et al.
Pubblicazione: (2025)
di: Li, Zhu, et al.
Pubblicazione: (2025)
Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
di: Li, Weiqin, et al.
Pubblicazione: (2024)
di: Li, Weiqin, et al.
Pubblicazione: (2024)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
di: Zhang, Zixing, et al.
Pubblicazione: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
di: Han, Bing, et al.
Pubblicazione: (2024)
di: Han, Bing, et al.
Pubblicazione: (2024)
On the Contribution of Lexical Features to Speech Emotion Recognition
di: Combei, David
Pubblicazione: (2025)
di: Combei, David
Pubblicazione: (2025)
CAMEO: Collection of Multilingual Emotional Speech Corpora
di: Christop, Iwona, et al.
Pubblicazione: (2025)
di: Christop, Iwona, et al.
Pubblicazione: (2025)
nEMO: Dataset of Emotional Speech in Polish
di: Christop, Iwona
Pubblicazione: (2024)
di: Christop, Iwona
Pubblicazione: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Next Tokens Denoising for Speech Synthesis
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
di: Liu, Yanqing, et al.
Pubblicazione: (2025)
A Cross-Corpus Speech Emotion Recognition Method Based on Supervised Contrastive Learning
di: minjie, Xiang
Pubblicazione: (2024)
di: minjie, Xiang
Pubblicazione: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
Improving Audio Codec-based Zero-Shot Text-to-Speech Synthesis with Multi-Modal Context and Large Language Model
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
di: Xue, Jinlong, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Generative Expressive Conversational Speech Synthesis
di: Liu, Rui, et al.
Pubblicazione: (2024)
di: Liu, Rui, et al.
Pubblicazione: (2024)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
di: Li, Yingting, et al.
Pubblicazione: (2024) -
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
di: Li, Yingting, et al.
Pubblicazione: (2024) -
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
di: Li, Xiang, et al.
Pubblicazione: (2024) -
Improving Text-To-Audio Models with Synthetic Captions
di: Kong, Zhifeng, et al.
Pubblicazione: (2024) -
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
di: Hung, Chia-Yu, et al.
Pubblicazione: (2024)