Optimizing Multilingual Text-To-Speech with Accents & Emotions
Fuente:
arXiv
Saved in:
| Main Authors: | Pawar, Pranav, Dwivedi, Akshansh, Boricha, Jenish, Gohil, Himanshu, Dubey, Aditya |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Human Feedback Driven Dynamic Speech Emotion Recognition
by: Fedorov, Ilya, et al.
Published: (2025)
by: Fedorov, Ilya, et al.
Published: (2025)
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
by: Sankey-Olsen, Cuno, et al.
Published: (2025)
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
by: Tiwari, Upasana, et al.
Published: (2025)
by: Tiwari, Upasana, et al.
Published: (2025)
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
by: Ali, Hasmot, et al.
Published: (2024)
by: Ali, Hasmot, et al.
Published: (2024)
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
by: Yuan, Kuang, et al.
Published: (2025)
by: Yuan, Kuang, et al.
Published: (2025)
VoXtream: Full-Stream Text-to-Speech with Extremely Low Latency
by: Torgashov, Nikita, et al.
Published: (2025)
by: Torgashov, Nikita, et al.
Published: (2025)
Quality Audio Prototyping: a prototype system for unified sound retrieval and procedural generation
by: Garcia, Nelly, et al.
Published: (2026)
by: Garcia, Nelly, et al.
Published: (2026)
Towards LLM-Empowered Fine-Grained Speech Descriptors for Explainable Emotion Recognition
by: Chen, Youjun, et al.
Published: (2025)
by: Chen, Youjun, et al.
Published: (2025)
VoiceX: A Text-To-Speech Framework for Custom Voices
by: Mertes, Silvan, et al.
Published: (2024)
by: Mertes, Silvan, et al.
Published: (2024)
AIx Speed: Playback Speed Optimization Using Listening Comprehension of Speech Recognition Models
by: Kawamura, Kazuki, et al.
Published: (2024)
by: Kawamura, Kazuki, et al.
Published: (2024)
Adaptation and Optimization of Automatic Speech Recognition (ASR) for the Maritime Domain in the Field of VHF Communication
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
by: Nakilcioglu, Emin Cagatay, et al.
Published: (2023)
Spontaneous Informal Speech Dataset for Punctuation Restoration
by: Liu, Xing Yi, et al.
Published: (2024)
by: Liu, Xing Yi, et al.
Published: (2024)
Personalized Speech Emotion Recognition in Human-Robot Interaction using Vision Transformers
by: Mishra, Ruchik, et al.
Published: (2024)
by: Mishra, Ruchik, et al.
Published: (2024)
Abjad-Kids: An Arabic Speech Classification Dataset for Primary Education
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
by: Snoubara, Abdul Aziz, et al.
Published: (2026)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Harnessing Smartwatch Microphone Sensors for Cough Detection and Classification
by: Jaiswal, Pranay, et al.
Published: (2024)
by: Jaiswal, Pranay, et al.
Published: (2024)
Improving AI-generated music with user-guided training
by: Singh, Vishwa Mohan, et al.
Published: (2025)
by: Singh, Vishwa Mohan, et al.
Published: (2025)
DOO-RE: A dataset of ambient sensors in a meeting room for activity recognition
by: Kim, Hyunju, et al.
Published: (2024)
by: Kim, Hyunju, et al.
Published: (2024)
Voice Passing : a Non-Binary Voice Gender Prediction System for evaluating Transgender voice transition
by: Doukhan, David, et al.
Published: (2024)
by: Doukhan, David, et al.
Published: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
by: Li, Yue, et al.
Published: (2024)
by: Li, Yue, et al.
Published: (2024)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
by: Zhao, Zhixian, et al.
Published: (2024)
by: Zhao, Zhixian, et al.
Published: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
by: Park, Seohyun, et al.
Published: (2025)
by: Park, Seohyun, et al.
Published: (2025)
Towards Decoding Brain Activity During Passive Listening of Speech
by: Fodor, Milán András, et al.
Published: (2024)
by: Fodor, Milán András, et al.
Published: (2024)
SACM: SEEG-Audio Contrastive Matching for Chinese Speech Decoding
by: Wang, Hongbin, et al.
Published: (2025)
by: Wang, Hongbin, et al.
Published: (2025)
Directional Source Separation for Robust Speech Recognition on Smart Glasses
by: Feng, Tiantian, et al.
Published: (2023)
by: Feng, Tiantian, et al.
Published: (2023)
Psychophysiology-aided Perceptually Fluent Speech Analysis of Children Who Stutter
by: Xiao, Yi, et al.
Published: (2022)
by: Xiao, Yi, et al.
Published: (2022)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
by: Khanday, Owais Mujtaba, et al.
Published: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
by: Nowrin, Sadia, et al.
Published: (2024)
by: Nowrin, Sadia, et al.
Published: (2024)
Layer-Wise Analysis of Self-Supervised Representations for Age and Gender Classification in Children's Speech
by: Sinha, Abhijit, et al.
Published: (2025)
by: Sinha, Abhijit, et al.
Published: (2025)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
by: Yu, Luca Jiang-Tao, et al.
Published: (2024)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
by: Dutta, Satwik, et al.
Published: (2025)
by: Dutta, Satwik, et al.
Published: (2025)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
by: Ma, Yong, et al.
Published: (2025)
by: Ma, Yong, et al.
Published: (2025)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
by: Liu, Ailin, et al.
Published: (2024)
by: Liu, Ailin, et al.
Published: (2024)
Clip-TTS: Contrastive Text-content and Mel-spectrogram, A High-Quality Text-to-Speech Method based on Contextual Semantic Understanding
by: Liu, Tianyun
Published: (2025)
by: Liu, Tianyun
Published: (2025)
STAA-Net: A Sparse and Transferable Adversarial Attack for Speech Emotion Recognition
by: Chang, Yi, et al.
Published: (2024)
by: Chang, Yi, et al.
Published: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
by: Melechovsky, Jan, et al.
Published: (2022)
by: Melechovsky, Jan, et al.
Published: (2022)
Audio Turing Test: Benchmarking the Human-likeness of Large Language Model-based Text-to-Speech Systems in Chinese
by: Wang, Xihuai, et al.
Published: (2025)
by: Wang, Xihuai, et al.
Published: (2025)
Emotion Detection Using Conditional Generative Adversarial Networks (cGAN): A Deep Learning Approach
by: Srivastava, Anushka
Published: (2025)
by: Srivastava, Anushka
Published: (2025)
VoiceFlow: Efficient Text-to-Speech with Rectified Flow Matching
by: Guo, Yiwei, et al.
Published: (2023)
by: Guo, Yiwei, et al.
Published: (2023)
Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital Avatars
by: NVIDIA, et al.
Published: (2025)
by: NVIDIA, et al.
Published: (2025)
Similar Items
-
Human Feedback Driven Dynamic Speech Emotion Recognition
by: Fedorov, Ilya, et al.
Published: (2025) -
Detecting COPD Through Speech Analysis: A Dataset of Danish Speech and Machine Learning Approach
by: Sankey-Olsen, Cuno, et al.
Published: (2025) -
Emotion-Disentangled Embedding Alignment for Noise-Robust and Cross-Corpus Speech Emotion Recognition
by: Tiwari, Upasana, et al.
Published: (2025) -
Artificial Neural Networks to Recognize Speakers Division from Continuous Bengali Speech
by: Ali, Hasmot, et al.
Published: (2024) -
SonicSieve: Bringing Directional Speech Extraction to Smartphones Using Acoustic Microstructures
by: Yuan, Kuang, et al.
Published: (2025)