EmoVoice: LLM-based Emotional Text-To-Speech Model with Freestyle Text Prompting
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Guanrou, Yang, Chen, Chen, Qian, Ma, Ziyang, Chen, Wenxi, Wang, Wen, Wang, Tianrui, Yang, Yifan, Niu, Zhikang, Liu, Wenrui, Yu, Fan, Du, Zhihao, Gao, Zhifu, Zhang, ShiLiang, Chen, Xie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
di: Yang, Guanrou, et al.
Pubblicazione: (2026)
di: Yang, Guanrou, et al.
Pubblicazione: (2026)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
di: Tu, Wenming, et al.
Pubblicazione: (2025)
di: Tu, Wenming, et al.
Pubblicazione: (2025)
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
di: Song, Yakun, et al.
Pubblicazione: (2024)
di: Song, Yakun, et al.
Pubblicazione: (2024)
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
di: Zheng, Qixi, et al.
Pubblicazione: (2025)
di: Zheng, Qixi, et al.
Pubblicazione: (2025)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
di: Yang, Yifan, et al.
Pubblicazione: (2026)
di: Yang, Yifan, et al.
Pubblicazione: (2026)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
di: Du, Chenpeng, et al.
Pubblicazione: (2024)
di: Du, Chenpeng, et al.
Pubblicazione: (2024)
CTC-Assisted LLM-Based Contextual ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
MaLa-ASR: Multimedia-Assisted LLM-Based ASR
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
di: Yang, Guanrou, et al.
Pubblicazione: (2024)
EmoBox: Multilingual Multi-corpus Speech Emotion Recognition Toolkit and Benchmark
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
di: Yang, Yifan, et al.
Pubblicazione: (2024)
di: Yang, Yifan, et al.
Pubblicazione: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
di: Abilbekov, Adal, et al.
Pubblicazione: (2024)
di: Abilbekov, Adal, et al.
Pubblicazione: (2024)
Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
di: Liu, Wenrui, et al.
Pubblicazione: (2025)
An Embarrassingly Simple Approach for LLM with Strong ASR Capacity
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
di: Ma, Ziyang, et al.
Pubblicazione: (2024)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
di: Yan, Ruiqi, et al.
Pubblicazione: (2025)
Position: Towards Responsible Evaluation for Text-to-Speech
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
di: Ma, Ziyang, et al.
Pubblicazione: (2023)
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
di: Choi, Jeongsoo, et al.
Pubblicazione: (2025)
CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
di: Du, Zhihao, et al.
Pubblicazione: (2024)
di: Du, Zhihao, et al.
Pubblicazione: (2024)
SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training
di: Chen, Wenxi, et al.
Pubblicazione: (2024)
di: Chen, Wenxi, et al.
Pubblicazione: (2024)
CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
di: Du, Zhihao, et al.
Pubblicazione: (2025)
di: Du, Zhihao, et al.
Pubblicazione: (2025)
SAC: Neural Speech Codec with Semantic-Acoustic Dual-Stream Quantization
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
di: Chen, Wenxi, et al.
Pubblicazione: (2025)
EmoAttack: Utilizing Emotional Voice Conversion for Speech Backdoor Attacks on Deep Speech Classification Models
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
di: Yao, Wenhan, et al.
Pubblicazione: (2024)
DiSTAR: Diffusion over a Scalable Token Autoregressive Representation for Speech Generation
di: Song, Yakun, et al.
Pubblicazione: (2025)
di: Song, Yakun, et al.
Pubblicazione: (2025)
Towards Flow-Matching-based TTS without Classifier-Free Guidance
di: Liang, Yuzhe, et al.
Pubblicazione: (2025)
di: Liang, Yuzhe, et al.
Pubblicazione: (2025)
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
di: Chen, Yushen, et al.
Pubblicazione: (2024)
di: Chen, Yushen, et al.
Pubblicazione: (2024)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
Evaluating the Expressive Appropriateness of Speech in Rich Contexts
di: Wang, Tianrui, et al.
Pubblicazione: (2026)
di: Wang, Tianrui, et al.
Pubblicazione: (2026)
Emo-DPO: Controllable Emotional Speech Synthesis through Direct Preference Optimization
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
di: Gao, Xiaoxue, et al.
Pubblicazione: (2024)
EmoQ: Speech Emotion Recognition via Speech-Aware Q-Former and Large Language Model
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
di: Yang, Yiqing, et al.
Pubblicazione: (2025)
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
Semantic-VAE: Semantic-Alignment Latent Representation for Better Speech Synthesis
di: Niu, Zhikang, et al.
Pubblicazione: (2025)
di: Niu, Zhikang, et al.
Pubblicazione: (2025)
EmoFormer: A Text-Independent Speech Emotion Recognition using a Hybrid Transformer-CNN model
di: Hasan, Rashedul, et al.
Pubblicazione: (2025)
di: Hasan, Rashedul, et al.
Pubblicazione: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
di: Guo, Yiwei, et al.
Pubblicazione: (2023)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
di: Zhou, Li, et al.
Pubblicazione: (2026)
di: Zhou, Li, et al.
Pubblicazione: (2026)
Drop the beat! Freestyler for Accompaniment Conditioned Rapping Voice Generation
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
di: Ning, Ziqian, et al.
Pubblicazione: (2024)
EmoSphere++: Emotion-Controllable Zero-Shot Text-to-Speech via Emotion-Adaptive Spherical Vector
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
di: Cho, Deok-Hyeon, et al.
Pubblicazione: (2024)
Voice Attribute Editing with Text Prompt
di: Sheng, Zhengyan, et al.
Pubblicazione: (2024)
di: Sheng, Zhengyan, et al.
Pubblicazione: (2024)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
di: Chen, Zhengyang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
WavCube: Unifying Speech Representation for Understanding and Generation via Semantic-Acoustic Joint Modeling
di: Yang, Guanrou, et al.
Pubblicazione: (2026) -
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025) -
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models
di: Tu, Wenming, et al.
Pubblicazione: (2025) -
TacoLM: GaTed Attention Equipped Codec Language Model are Efficient Zero-Shot Text to Speech Synthesizers
di: Song, Yakun, et al.
Pubblicazione: (2024) -
Accelerating Flow-Matching-Based Text-to-Speech via Empirically Pruned Step Sampling
di: Zheng, Qixi, et al.
Pubblicazione: (2025)