Toward Natural Emotional Text-To-Speech System with Fine-Grained Non-Verbal Expression Control
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Wangzixi, Atmaja, Bagus Tris, Sakti, Sakriani |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Uncertainty-Based Ensemble Learning For Speech Classification
di: Atmaja, Bagus Tris, et al.
Pubblicazione: (2024)
di: Atmaja, Bagus Tris, et al.
Pubblicazione: (2024)
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
di: Zhou, Wangzixi, et al.
Pubblicazione: (2026)
di: Zhou, Wangzixi, et al.
Pubblicazione: (2026)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
di: Hirano, Yuta, et al.
Pubblicazione: (2025)
di: Hirano, Yuta, et al.
Pubblicazione: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
di: Girish, et al.
Pubblicazione: (2026)
di: Girish, et al.
Pubblicazione: (2026)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
di: Adila, Aulia, et al.
Pubblicazione: (2024)
di: Adila, Aulia, et al.
Pubblicazione: (2024)
Learning Marmoset Vocal Patterns with a Masked Autoencoder for Robust Call Segmentation, Classification, and Caller Identification
di: Wu, Bin, et al.
Pubblicazione: (2024)
di: Wu, Bin, et al.
Pubblicazione: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
di: Xin, Detai, et al.
Pubblicazione: (2023)
di: Xin, Detai, et al.
Pubblicazione: (2023)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
di: Shen, Siyuan, et al.
Pubblicazione: (2024)
Fine-Grained Quantitative Emotion Editing for Speech Generation
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Indonesian-English Code-Switching Speech Synthesizer Utilizing Multilingual STEN-TTS and Bert LID
di: Handoyo, Ahmad Alfani, et al.
Pubblicazione: (2024)
di: Handoyo, Ahmad Alfani, et al.
Pubblicazione: (2024)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
Continual Learning in Machine Speech Chain Using Gradient Episodic Memory
di: Tyndall, Geoffrey, et al.
Pubblicazione: (2024)
di: Tyndall, Geoffrey, et al.
Pubblicazione: (2024)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Enhancing Emotional Text-to-Speech Controllability with Natural Language Guidance through Contrastive Learning and Diffusion Models
di: Jing, Xin, et al.
Pubblicazione: (2024)
di: Jing, Xin, et al.
Pubblicazione: (2024)
Towards Fine-Grained and Multi-Granular Contrastive Language-Speech Pre-training
di: Yang, Yifan, et al.
Pubblicazione: (2026)
di: Yang, Yifan, et al.
Pubblicazione: (2026)
Controlling Emotion in Text-to-Speech with Natural Language Prompts
di: Bott, Thomas, et al.
Pubblicazione: (2024)
di: Bott, Thomas, et al.
Pubblicazione: (2024)
FlexSpeech: Towards Stable, Controllable and Expressive Text-to-Speech
di: Ma, Linhan, et al.
Pubblicazione: (2025)
di: Ma, Linhan, et al.
Pubblicazione: (2025)
NAIST Simultaneous Speech Translation System for IWSLT 2024
di: Ko, Yuka, et al.
Pubblicazione: (2024)
di: Ko, Yuka, et al.
Pubblicazione: (2024)
A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding
di: Ye, Runchuan, et al.
Pubblicazione: (2025)
di: Ye, Runchuan, et al.
Pubblicazione: (2025)
AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
di: Qi, Tianhua, et al.
Pubblicazione: (2026)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2025)
di: Inoue, Sho, et al.
Pubblicazione: (2025)
Towards Fine-Grained Multi-Dimensional Speech Understanding: Data Pipeline, Benchmark, and Model
di: Li, Guojian, et al.
Pubblicazione: (2026)
di: Li, Guojian, et al.
Pubblicazione: (2026)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
di: Li, Haoyang, et al.
Pubblicazione: (2026)
di: Li, Haoyang, et al.
Pubblicazione: (2026)
Intelligibility of Text-to-Speech Systems for Mathematical Expressions
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2025)
di: Roychowdhury, Sujoy, et al.
Pubblicazione: (2025)
Hierarchical Control of Emotion Rendering in Speech Synthesis
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
di: Chen, Youjun, et al.
Pubblicazione: (2025)
di: Chen, Youjun, et al.
Pubblicazione: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
di: Wang, Guansu, et al.
Pubblicazione: (2025)
di: Wang, Guansu, et al.
Pubblicazione: (2025)
Are Mamba-based Audio Foundation Models the Best Fit for Non-Verbal Emotion Recognition?
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
di: Akhtar, Mohd Mujtaba, et al.
Pubblicazione: (2025)
Learning Fine-Grained Controllability on Speech Generation via Efficient Fine-Tuning
di: Chien, Chung-Ming, et al.
Pubblicazione: (2024)
di: Chien, Chung-Ming, et al.
Pubblicazione: (2024)
Fine-Grained and Interpretable Neural Speech Editing
di: Morrison, Max, et al.
Pubblicazione: (2024)
di: Morrison, Max, et al.
Pubblicazione: (2024)
Position: Towards Responsible Evaluation for Text-to-Speech
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Textless and Non-Parallel Speech-to-Speech Emotion Style Transfer
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
di: Dutta, Soumya, et al.
Pubblicazione: (2025)
TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
di: Zhou, Junzuo, et al.
Pubblicazione: (2024)
di: Zhou, Junzuo, et al.
Pubblicazione: (2024)
KazEmoTTS: A Dataset for Kazakh Emotional Text-to-Speech Synthesis
di: Abilbekov, Adal, et al.
Pubblicazione: (2024)
di: Abilbekov, Adal, et al.
Pubblicazione: (2024)
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
di: Zhang, Xu, et al.
Pubblicazione: (2026)
di: Zhang, Xu, et al.
Pubblicazione: (2026)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
di: Yao, Jixun, et al.
Pubblicazione: (2025)
di: Yao, Jixun, et al.
Pubblicazione: (2025)
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech
di: Zhou, Enting, et al.
Pubblicazione: (2023)
di: Zhou, Enting, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Uncertainty-Based Ensemble Learning For Speech Classification
di: Atmaja, Bagus Tris, et al.
Pubblicazione: (2024) -
WaveNeXt 2: ConvNeXt-Based Fast Neural Vocoders With Residual Denoising and Sub-Modeling for GAN and Diffusion Models
di: Zhou, Wangzixi, et al.
Pubblicazione: (2026) -
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
di: Hirano, Yuta, et al.
Pubblicazione: (2025) -
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
di: Rossenbach, Nick, et al.
Pubblicazione: (2024) -
Prosody as Supervision: Bridging the Non-Verbal--Verbal for Multilingual Speech Emotion Recognition
di: Girish, et al.
Pubblicazione: (2026)