Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
Fuente:
arXiv
Saved in:
| Main Authors: | Melechovsky, Jan, Mehrish, Ambuj, Sisman, Berrak, Herremans, Dorien |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
by: Melechovsky, Jan, et al.
Published: (2022)
by: Melechovsky, Jan, et al.
Published: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
by: Lam, Perry, et al.
Published: (2022)
by: Lam, Perry, et al.
Published: (2022)
MidiCaps: A large-scale MIDI dataset with text captions
by: Melechovsky, Jan, et al.
Published: (2024)
by: Melechovsky, Jan, et al.
Published: (2024)
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
by: Mahapatra, Aurosweta, et al.
Published: (2025)
by: Mahapatra, Aurosweta, et al.
Published: (2025)
Discrete Unit based Masking for Improving Disentanglement in Voice Conversion
by: Lee, Philip H., et al.
Published: (2024)
by: Lee, Philip H., et al.
Published: (2024)
PRESENT: Zero-Shot Text-to-Prosody Control
by: Lam, Perry, et al.
Published: (2024)
by: Lam, Perry, et al.
Published: (2024)
SonicMaster: Towards Controllable All-in-One Music Restoration and Mastering
by: Melechovsky, Jan, et al.
Published: (2025)
by: Melechovsky, Jan, et al.
Published: (2025)
Style Mixture of Experts for Expressive Text-To-Speech Synthesis
by: Jawaid, Ahad, et al.
Published: (2024)
by: Jawaid, Ahad, et al.
Published: (2024)
Revealing Emotional Clusters in Speaker Embeddings: A Contrastive Learning Strategy for Speech Emotion Recognition
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
by: Li, Yingting, et al.
Published: (2024)
by: Li, Yingting, et al.
Published: (2024)
Enhancing Speech Emotion Recognition Through Differentiable Architecture Search
by: Rajapakshe, Thejan, et al.
Published: (2023)
by: Rajapakshe, Thejan, et al.
Published: (2023)
Objective Evaluation of Prosody and Intelligibility in Speech Synthesis via Conditional Prediction of Discrete Tokens
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
by: Ulgen, Ismail Rasim, et al.
Published: (2025)
emoDARTS: Joint Optimisation of CNN & Sequential Neural Network Architectures for Superior Speech Emotion Recognition
by: Rajapakshe, Thejan, et al.
Published: (2024)
by: Rajapakshe, Thejan, et al.
Published: (2024)
Leveraging LLM Embeddings for Cross Dataset Label Alignment and Zero Shot Music Emotion Prediction
by: Liu, Renhang, et al.
Published: (2024)
by: Liu, Renhang, et al.
Published: (2024)
MelodySim: Measuring Melody-aware Music Similarity for Plagiarism Detection
by: Lu, Tongyu, et al.
Published: (2025)
by: Lu, Tongyu, et al.
Published: (2025)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
by: Inoue, Sho, et al.
Published: (2024)
by: Inoue, Sho, et al.
Published: (2024)
NaturalVoices: A Large-Scale, Spontaneous and Emotional Podcast Dataset for Voice Conversion
by: Du, Zongyang, et al.
Published: (2025)
by: Du, Zongyang, et al.
Published: (2025)
DiffAnon: Diffusion-based Prosody Control for Voice Anonymization
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
by: Ulgen, Ismail Rasim, et al.
Published: (2026)
Leveraging Parameter-Efficient Transfer Learning for Multi-Lingual Text-to-Speech Adaptation
by: Li, Yingting, et al.
Published: (2024)
by: Li, Yingting, et al.
Published: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
by: Zhang, Shaozuo, et al.
Published: (2025)
by: Zhang, Shaozuo, et al.
Published: (2025)
Can Emotion Fool Anti-spoofing?
by: Mahapatra, Aurosweta, et al.
Published: (2025)
by: Mahapatra, Aurosweta, et al.
Published: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
by: Zhou, Xuehao, et al.
Published: (2024)
by: Zhou, Xuehao, et al.
Published: (2024)
Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model
by: Du, Zongyang, et al.
Published: (2024)
by: Du, Zongyang, et al.
Published: (2024)
Versatile audio-visual learning for emotion recognition
by: Goncalves, Lucas, et al.
Published: (2023)
by: Goncalves, Lucas, et al.
Published: (2023)
Adapting Automatic Speech Recognition for Accented Air Traffic Control Communications
by: Wee, Marcus Yu Zhe, et al.
Published: (2025)
by: Wee, Marcus Yu Zhe, et al.
Published: (2025)
Improving Text-To-Audio Models with Synthetic Captions
by: Kong, Zhifeng, et al.
Published: (2024)
by: Kong, Zhifeng, et al.
Published: (2024)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
by: Pawar, Pranav, et al.
Published: (2025)
by: Pawar, Pranav, et al.
Published: (2025)
Are We There Yet? A Brief Survey of Music Emotion Prediction Datasets, Models and Outstanding Challenges
by: Kang, Jaeyong, et al.
Published: (2024)
by: Kang, Jaeyong, et al.
Published: (2024)
Towards Unified Music Emotion Recognition across Dimensional and Categorical Models
by: Kang, Jaeyong, et al.
Published: (2025)
by: Kang, Jaeyong, et al.
Published: (2025)
Aligning Generative Music AI with Human Preferences: Methods and Challenges
by: Herremans, Dorien, et al.
Published: (2025)
by: Herremans, Dorien, et al.
Published: (2025)
GE2E-AC: Generalized End-to-End Loss Training for Accent Classification
by: Watanabe, Chihiro, et al.
Published: (2024)
by: Watanabe, Chihiro, et al.
Published: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
by: Cheng, Zhuangfei, et al.
Published: (2025)
by: Cheng, Zhuangfei, et al.
Published: (2025)
Sequence-Level Unsupervised Training in Speech Recognition: A Theoretical Study
by: Yang, Zijian, et al.
Published: (2026)
by: Yang, Zijian, et al.
Published: (2026)
CM-TTS: Enhancing Real Time Text-to-Speech Synthesis Efficiency through Weighted Samplers and Consistency Models
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Adversarial Training of Denoising Diffusion Model Using Dual Discriminators for High-Fidelity Multi-Speaker TTS
by: Ko, Myeongjin, et al.
Published: (2023)
by: Ko, Myeongjin, et al.
Published: (2023)
EM-TTS: Efficiently Trained Low-Resource Mongolian Lightweight Text-to-Speech
by: Liang, Ziqi, et al.
Published: (2024)
by: Liang, Ziqi, et al.
Published: (2024)
Text-to-Speech for Unseen Speakers via Low-Complexity Discrete Unit-Based Frame Selection
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
by: Ulgen, Ismail Rasim, et al.
Published: (2024)
Multi-modal Adversarial Training for Zero-Shot Voice Cloning
by: Janiczek, John, et al.
Published: (2024)
by: Janiczek, John, et al.
Published: (2024)
Universal Speech Content Factorization
by: Xinyuan, Henry Li, et al.
Published: (2026)
by: Xinyuan, Henry Li, et al.
Published: (2026)
Similar Items
-
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
by: Melechovsky, Jan, et al.
Published: (2022) -
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
by: Melechovsky, Jan, et al.
Published: (2024) -
SNIPER Training: Single-Shot Sparse Training for Text-to-Speech
by: Lam, Perry, et al.
Published: (2022) -
MidiCaps: A large-scale MIDI dataset with text captions
by: Melechovsky, Jan, et al.
Published: (2024) -
HuLA: Prosody-Aware Anti-Spoofing with Multi-Task Learning for Expressive and Emotional Synthetic Speech
by: Mahapatra, Aurosweta, et al.
Published: (2025)