Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Inoue, Sho, Zhou, Kun, Wang, Shuai, Li, Haizhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
von: Inoue, Sho, et al.
Veröffentlicht: (2025)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
von: Liu, Zhijun, et al.
Veröffentlicht: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
Context-Aware Two-Step Training Scheme for Domain Invariant Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
von: Wang, Tianrui, et al.
Veröffentlicht: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
von: Qi, Tianhua, et al.
Veröffentlicht: (2026)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
von: Jing, Xin, et al.
Veröffentlicht: (2024)
von: Jing, Xin, et al.
Veröffentlicht: (2024)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
von: Nishimura, Yuto, et al.
Veröffentlicht: (2024)
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
von: Li, Sirui, et al.
Veröffentlicht: (2025)
von: Li, Sirui, et al.
Veröffentlicht: (2025)
Causal Self-supervised Pretrained Frontend with Predictive Code for Speech Separation
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
von: Wang, Wupeng, et al.
Veröffentlicht: (2025)
ED-TTS: Multi-Scale Emotion Modeling using Cross-Domain Emotion Diarization for Emotional Speech Synthesis
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
von: Tang, Haobin, et al.
Veröffentlicht: (2024)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
von: Hu, Yifan, et al.
Veröffentlicht: (2025)
Hierarchical Decoding for Discrete Speech Synthesis with Multi-Resolution Spoof Detection
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
Emotional Dimension Control in Language Model-Based Text-to-Speech: Spanning a Broad Spectrum of Human Emotions
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
von: Zhou, Kun, et al.
Veröffentlicht: (2024)
M-Vec: Matryoshka Speaker Embeddings with Flexible Dimensions
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
SoCodec: A Semantic-Ordered Multi-Stream Speech Codec for Efficient Language Model Based Text-to-Speech Synthesis
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
von: Guo, Haohan, et al.
Veröffentlicht: (2024)
Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
von: Liao, Shijia, et al.
Veröffentlicht: (2024)
EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
von: Zhou, Li, et al.
Veröffentlicht: (2026)
von: Zhou, Li, et al.
Veröffentlicht: (2026)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
von: Zhang, Shaozuo, et al.
Veröffentlicht: (2025)
MSceneSpeech: A Multi-Scene Speech Dataset For Expressive Speech Synthesis
von: Yang, Qian, et al.
Veröffentlicht: (2024)
von: Yang, Qian, et al.
Veröffentlicht: (2024)
EME-TTS: Unlocking the Emphasis and Emotion Link in Speech Synthesis
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
von: Li, Haoxun, et al.
Veröffentlicht: (2025)
MM-TTS: Multi-modal Prompt based Style Transfer for Expressive Text-to-Speech Synthesis
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
von: Guan, Wenhao, et al.
Veröffentlicht: (2023)
EELE: Exploring Efficient and Extensible LoRA Integration in Emotional Text-to-Speech
von: Qi, Xin, et al.
Veröffentlicht: (2024)
von: Qi, Xin, et al.
Veröffentlicht: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
von: Zhang, Ke, et al.
Veröffentlicht: (2024)
Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
von: Du, Chenpeng, et al.
Veröffentlicht: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
von: Zhang, Xu, et al.
Veröffentlicht: (2026)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
von: Shen, Siyuan, et al.
Veröffentlicht: (2024)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
von: Jiang, Yuepeng, et al.
Veröffentlicht: (2024)
PhiNet: Speaker Verification with Phonetic Interpretability
von: Ma, Yi, et al.
Veröffentlicht: (2026)
von: Ma, Yi, et al.
Veröffentlicht: (2026)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
von: Zhang, Leying, et al.
Veröffentlicht: (2025)
DAIEN-TTS: Disentangled Audio Infilling for Environment-Aware Text-to-Speech Synthesis
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
von: Lu, Ye-Xin, et al.
Veröffentlicht: (2025)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
von: Lei, Shun, et al.
Veröffentlicht: (2023)
von: Lei, Shun, et al.
Veröffentlicht: (2023)
FluentEditor2: Text-based Speech Editing by Modeling Multi-Scale Acoustic and Prosody Consistency
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
RSET: Remapping-based Sorting Method for Emotion Transfer Speech Synthesis
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
Hierarchical Control of Emotion Rendering in Speech Synthesis
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
Fine-Grained Quantitative Emotion Editing for Speech Generation
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024) -
PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
von: Inoue, Sho, et al.
Veröffentlicht: (2025)