Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
Fuente:
arXiv
Guardado en:
| Autores principales: | Mehta, Shivam, Lameris, Harm, Punmiya, Rajiv, Beskow, Jonas, Székely, Éva, Henter, Gustav Eje |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Matcha-TTS: A fast TTS architecture with conditional flow matching
por: Mehta, Shivam, et al.
Publicado: (2023)
por: Mehta, Shivam, et al.
Publicado: (2023)
Unified speech and gesture synthesis using flow matching
por: Mehta, Shivam, et al.
Publicado: (2023)
por: Mehta, Shivam, et al.
Publicado: (2023)
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
por: Mehta, Shivam, et al.
Publicado: (2024)
por: Mehta, Shivam, et al.
Publicado: (2024)
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
por: Deichler, Anna, et al.
Publicado: (2025)
por: Deichler, Anna, et al.
Publicado: (2025)
IMUVIE: Pickup Timeline Action Localization via Motion Movies
por: Clapham, John, et al.
Publicado: (2024)
por: Clapham, John, et al.
Publicado: (2024)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
por: Semenov, Andrei, et al.
Publicado: (2024)
por: Semenov, Andrei, et al.
Publicado: (2024)
Make Some Noise: Towards LLM audio reasoning and generation using sound tokens
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
Uncovering Population PK Covariates from VAE-Generated Latent Spaces
por: Perazzolo, Diego, et al.
Publicado: (2025)
por: Perazzolo, Diego, et al.
Publicado: (2025)
Adaptive Riemannian Graph Neural Networks
por: Wang, Xudong, et al.
Publicado: (2025)
por: Wang, Xudong, et al.
Publicado: (2025)
Learnable Kernel Density Estimation for Graphs and Its Application to Graph-Level Anomaly Detection
por: Wang, Xudong, et al.
Publicado: (2025)
por: Wang, Xudong, et al.
Publicado: (2025)
A Spatio-Temporal Deep Learning Approach For High-Resolution Gridded Monsoon Prediction
por: Borah, Parashjyoti, et al.
Publicado: (2026)
por: Borah, Parashjyoti, et al.
Publicado: (2026)
Improving Omics-Based Classification: The Role of Feature Selection and Synthetic Data Generation
por: Perazzolo, Diego, et al.
Publicado: (2025)
por: Perazzolo, Diego, et al.
Publicado: (2025)
An accurate and revised version of optical character recognition-based speech synthesis using LabVIEW
por: Mehta, Prateek, et al.
Publicado: (2025)
por: Mehta, Prateek, et al.
Publicado: (2025)
Towards Context-Aware Human-like Pointing Gestures with RL Motion Imitation
por: Deichler, Anna, et al.
Publicado: (2025)
por: Deichler, Anna, et al.
Publicado: (2025)
Adversarially Probing Cross-Family Sound Symbolism in 27 Languages
por: Sharma, Anika, et al.
Publicado: (2025)
por: Sharma, Anika, et al.
Publicado: (2025)
Regularisation in neural networks: a survey and empirical analysis of approaches
por: Opperman, Christiaan P., et al.
Publicado: (2026)
por: Opperman, Christiaan P., et al.
Publicado: (2026)
Alternative Local Discriminant Bases Using Empirical Expectation and Variance Estimation
por: Fossgaard, Eirik
Publicado: (1999)
por: Fossgaard, Eirik
Publicado: (1999)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
por: Deichler, Anna, et al.
Publicado: (2026)
por: Deichler, Anna, et al.
Publicado: (2026)
Balancing the Scales: A Comprehensive Study on Tackling Class Imbalance in Binary Classification
por: Abdelhamid, Mohamed, et al.
Publicado: (2024)
por: Abdelhamid, Mohamed, et al.
Publicado: (2024)
Transformers Meet Relational Databases
por: Peleška, Jakub, et al.
Publicado: (2024)
por: Peleška, Jakub, et al.
Publicado: (2024)
Discriminative Subspace Emersion from learning feature relevances across different populations
por: Canducci, Marco, et al.
Publicado: (2025)
por: Canducci, Marco, et al.
Publicado: (2025)
SpATr: MoCap 3D Human Action Recognition based on Spiral Auto-encoder and Transformer Network
por: Bouzid, Hamza, et al.
Publicado: (2023)
por: Bouzid, Hamza, et al.
Publicado: (2023)
ScaleMAP: Preserving Local Density and Neighborhood Structure in Low-Dimensional Embeddings
por: Poorna, Rajas, et al.
Publicado: (2026)
por: Poorna, Rajas, et al.
Publicado: (2026)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
por: Lapin, Mykyta, et al.
Publicado: (2025)
por: Lapin, Mykyta, et al.
Publicado: (2025)
CASE: Contrastive Activation for Saliency Estimation
por: Williamson, Dane, et al.
Publicado: (2025)
por: Williamson, Dane, et al.
Publicado: (2025)
Instruction and Solution Probabilities as Heuristics for Inductive Programming
por: McDaid, Edward, et al.
Publicado: (2025)
por: McDaid, Edward, et al.
Publicado: (2025)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
por: Du, Guanchen, et al.
Publicado: (2025)
por: Du, Guanchen, et al.
Publicado: (2025)
Devanagari Handwritten Character Recognition using Convolutional Neural Network
por: Mehta, Diksha, et al.
Publicado: (2025)
por: Mehta, Diksha, et al.
Publicado: (2025)
Continuous Latent Contexts Enable Efficient Online Learning in Transformers
por: Anand, Emile, et al.
Publicado: (2026)
por: Anand, Emile, et al.
Publicado: (2026)
PRACH Preamble Detection as a Multi-Class Classification Problem: A Machine Learning Approach Using SVM
por: Ferenc, F., et al.
Publicado: (2025)
por: Ferenc, F., et al.
Publicado: (2025)
What Questions Should Robots Be Able to Answer? A Dataset of User Questions for Explainable Robotics
por: Wachowiak, Lennart, et al.
Publicado: (2025)
por: Wachowiak, Lennart, et al.
Publicado: (2025)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
por: Martin, Michael R., et al.
Publicado: (2025)
por: Martin, Michael R., et al.
Publicado: (2025)
SemAlignVC: Enhancing zero-shot timbre conversion using semantic alignment
por: Mehta, Shivam, et al.
Publicado: (2025)
por: Mehta, Shivam, et al.
Publicado: (2025)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
por: Zaragozá, Lucía Gómez, et al.
Publicado: (2024)
Cryptogenic stroke and migraine: using probabilistic independence and machine learning to uncover latent sources of disease from the electronic health record
por: Betts, Joshua W., et al.
Publicado: (2025)
por: Betts, Joshua W., et al.
Publicado: (2025)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
por: Perera, Amal S., et al.
Publicado: (2025)
por: Perera, Amal S., et al.
Publicado: (2025)
Classification of descriptions and summary using multiple passes of statistical and natural language toolkits
por: Banthia, Saumya, et al.
Publicado: (2020)
por: Banthia, Saumya, et al.
Publicado: (2020)
Precision at Scale: Domain-Specific Datasets On-Demand
por: Rodríguez-de-Vera, Jesús M, et al.
Publicado: (2024)
por: Rodríguez-de-Vera, Jesús M, et al.
Publicado: (2024)
Grounded Gesture Generation: Language, Motion, and Space
por: Deichler, Anna, et al.
Publicado: (2025)
por: Deichler, Anna, et al.
Publicado: (2025)
Named entity recognition for Serbian legal documents: Design, methodology and dataset development
por: Kalušev, Vladimir, et al.
Publicado: (2025)
por: Kalušev, Vladimir, et al.
Publicado: (2025)
Ejemplares similares
-
Matcha-TTS: A fast TTS architecture with conditional flow matching
por: Mehta, Shivam, et al.
Publicado: (2023) -
Unified speech and gesture synthesis using flow matching
por: Mehta, Shivam, et al.
Publicado: (2023) -
Fake it to make it: Using synthetic data to remedy the data shortage in joint multimodal speech-and-gesture synthesis
por: Mehta, Shivam, et al.
Publicado: (2024) -
Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
por: Deichler, Anna, et al.
Publicado: (2025) -
IMUVIE: Pickup Timeline Action Localization via Motion Movies
por: Clapham, John, et al.
Publicado: (2024)