NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech
Fuente:
arXiv
Guardado en:
| Autores principales: | Borisov, Maksim, Spirin, Egor, Diatlova, Daria |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adapting WavLM for Speech Emotion Recognition
por: Diatlova, Daria, et al.
Publicado: (2024)
por: Diatlova, Daria, et al.
Publicado: (2024)
NonverbalTTS
por: Anonymous
Publicado: (2025)
por: Anonymous
Publicado: (2025)
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
por: Cho, Deok-Hyeon, et al.
Publicado: (2026)
por: Cho, Deok-Hyeon, et al.
Publicado: (2026)
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
por: Diatlova, Daria, et al.
Publicado: (2025)
por: Diatlova, Daria, et al.
Publicado: (2025)
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
por: Ni, Qinke, et al.
Publicado: (2026)
por: Ni, Qinke, et al.
Publicado: (2026)
METR: Image Watermarking with Large Number of Unique Messages
por: Varlamov, Alexander, et al.
Publicado: (2024)
por: Varlamov, Alexander, et al.
Publicado: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
por: Xin, Detai, et al.
Publicado: (2023)
por: Xin, Detai, et al.
Publicado: (2023)
LibriTTS-P: A Corpus with Speaking Style and Speaker Identity Prompts for Text-to-Speech and Style Captioning
por: Kawamura, Masaya, et al.
Publicado: (2024)
por: Kawamura, Masaya, et al.
Publicado: (2024)
TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models
por: Ji, Shengpeng, et al.
Publicado: (2023)
por: Ji, Shengpeng, et al.
Publicado: (2023)
Task Vector in TTS: Toward Emotionally Expressive Dialectal Speech Synthesis
por: Feng, Pengchao, et al.
Publicado: (2025)
por: Feng, Pengchao, et al.
Publicado: (2025)
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
por: Glazer, Neta, et al.
Publicado: (2025)
por: Glazer, Neta, et al.
Publicado: (2025)
SupertonicTTS: Towards Highly Efficient and Streamlined Text-to-Speech System
por: Kim, Hyeongju, et al.
Publicado: (2025)
por: Kim, Hyeongju, et al.
Publicado: (2025)
MNV-17: A High-Quality Performative Mandarin Dataset for Nonverbal Vocalization Recognition in Speech
por: Mai, Jialong, et al.
Publicado: (2025)
por: Mai, Jialong, et al.
Publicado: (2025)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
por: Liang, Ziqi, et al.
Publicado: (2024)
por: Liang, Ziqi, et al.
Publicado: (2024)
More Similar than Dissimilar: Modeling Annotators for Cross-Corpus Speech Emotion Recognition
por: Tavernor, James, et al.
Publicado: (2025)
por: Tavernor, James, et al.
Publicado: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
por: Li, Yingting, et al.
Publicado: (2024)
por: Li, Yingting, et al.
Publicado: (2024)
ASTRA: Aligning Speech and Text Representations for Asr without Sampling
por: Gaur, Neeraj, et al.
Publicado: (2024)
por: Gaur, Neeraj, et al.
Publicado: (2024)
JaCappella Corpus: A Japanese a Cappella Vocal Ensemble Corpus
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
por: Nakamura, Tomohiko, et al.
Publicado: (2022)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
por: Pawar, Pranav, et al.
Publicado: (2025)
por: Pawar, Pranav, et al.
Publicado: (2025)
BitTTS: Highly Compact Text-to-Speech Using 1.58-bit Quantization and Weight Indexing
por: Kawamura, Masaya, et al.
Publicado: (2025)
por: Kawamura, Masaya, et al.
Publicado: (2025)
IndicVoices-R: Unlocking a Massive Multilingual Multi-speaker Speech Corpus for Scaling Indian TTS
por: Sankar, Ashwin, et al.
Publicado: (2024)
por: Sankar, Ashwin, et al.
Publicado: (2024)
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving
por: Shankar, Bhavani, et al.
Publicado: (2024)
por: Shankar, Bhavani, et al.
Publicado: (2024)
CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
por: Wang, Siyi, et al.
Publicado: (2026)
por: Wang, Siyi, et al.
Publicado: (2026)
Mouth Articulation-Based Anchoring for Improved Cross-Corpus Speech Emotion Recognition
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
por: Upadhyay, Shreya G., et al.
Publicado: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
por: Lee, Keon, et al.
Publicado: (2024)
por: Lee, Keon, et al.
Publicado: (2024)
PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
por: Pahar, Madhurananda, et al.
Publicado: (2026)
por: Pahar, Madhurananda, et al.
Publicado: (2026)
Sample-Efficient Diffusion for Text-To-Speech Synthesis
por: Lovelace, Justin, et al.
Publicado: (2024)
por: Lovelace, Justin, et al.
Publicado: (2024)
Leveraging Speech PTM, Text LLM, and Emotional TTS for Speech Emotion Recognition
por: Ma, Ziyang, et al.
Publicado: (2023)
por: Ma, Ziyang, et al.
Publicado: (2023)
TED-TTS: Training-Free Intra-Utterance Emotion and Duration Control for Text-to-Speech Synthesis
por: Liang, Qifan, et al.
Publicado: (2026)
por: Liang, Qifan, et al.
Publicado: (2026)
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
por: Kumar, Sahil, et al.
Publicado: (2026)
por: Kumar, Sahil, et al.
Publicado: (2026)
A Dataset for Automatic Vocal Mode Classification
por: Hinrichs, Reemt, et al.
Publicado: (2026)
por: Hinrichs, Reemt, et al.
Publicado: (2026)
NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations
por: Liao, Huan, et al.
Publicado: (2025)
por: Liao, Huan, et al.
Publicado: (2025)
STTATTS: Unified Speech-To-Text And Text-To-Speech Model
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
por: Toyin, Hawau Olamide, et al.
Publicado: (2024)
HiFi-Stream: Streaming Speech Enhancement with Generative Adversarial Networks
por: Dmitrieva, Ekaterina, et al.
Publicado: (2025)
por: Dmitrieva, Ekaterina, et al.
Publicado: (2025)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
por: Liu, Tianyun
Publicado: (2025)
por: Liu, Tianyun
Publicado: (2025)
Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
por: Leem, Seong-Gyun, et al.
Publicado: (2024)
por: Leem, Seong-Gyun, et al.
Publicado: (2024)
Speech Emotion Recognition with Phonation Excitation Information and Articulatory Kinematics
por: Zhang, Ziqian, et al.
Publicado: (2025)
por: Zhang, Ziqian, et al.
Publicado: (2025)
Causal Prosody Mediation for Text-to-Speech:Counterfactual Training of Duration, Pitch, and Energy in FastSpeech2
por: Mohanty, Suvendu Sekhar
Publicado: (2026)
por: Mohanty, Suvendu Sekhar
Publicado: (2026)
Improving Speech Emotion Recognition with Mutual Information Regularized Generative Model
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
por: Ahn, Chung-Soo, et al.
Publicado: (2025)
TaDiCodec: Text-aware Diffusion Speech Tokenizer for Speech Language Modeling
por: Wang, Yuancheng, et al.
Publicado: (2025)
por: Wang, Yuancheng, et al.
Publicado: (2025)
Ejemplares similares
-
Adapting WavLM for Speech Emotion Recognition
por: Diatlova, Daria, et al.
Publicado: (2024) -
NonverbalTTS
por: Anonymous
Publicado: (2025) -
Affectron: Emotional Speech Synthesis with Affective and Contextually Aligned Nonverbal Vocalizations
por: Cho, Deok-Hyeon, et al.
Publicado: (2026) -
VARAN: Variational Inference for Self-Supervised Speech Models Fine-Tuning on Downstream Tasks
por: Diatlova, Daria, et al.
Publicado: (2025) -
NV-Bench: Benchmark of Nonverbal Vocalization Synthesis for Expressive Text-to-Speech Generation
por: Ni, Qinke, et al.
Publicado: (2026)