Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhou, Xuehao, Zhang, Mingyang, Zhou, Yi, Wu, Zhizheng, Li, Haizhou |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2024)
por: Mu, Bingshen, et al.
Publicado: (2024)
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
por: Bai, Qibing, et al.
Publicado: (2026)
por: Bai, Qibing, et al.
Publicado: (2026)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
por: Cheng, Zhuangfei, et al.
Publicado: (2025)
Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2025)
por: Inoue, Sho, et al.
Publicado: (2025)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
por: Melechovsky, Jan, et al.
Publicado: (2024)
por: Melechovsky, Jan, et al.
Publicado: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
por: Wang, Wenbin, et al.
Publicado: (2024)
por: Wang, Wenbin, et al.
Publicado: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
por: Melechovsky, Jan, et al.
Publicado: (2022)
por: Melechovsky, Jan, et al.
Publicado: (2022)
Study on the Fairness of Speaker Verification Systems on Underrepresented Accents in English
por: Estevez, Mariel, et al.
Publicado: (2022)
por: Estevez, Mariel, et al.
Publicado: (2022)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2025)
por: Mu, Bingshen, et al.
Publicado: (2025)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
por: Bafna, Niyati, et al.
Publicado: (2025)
por: Bafna, Niyati, et al.
Publicado: (2025)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
por: Huang, Wen-Chin, et al.
Publicado: (2026)
por: Huang, Wen-Chin, et al.
Publicado: (2026)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
por: Bai, Qibing, et al.
Publicado: (2025)
por: Bai, Qibing, et al.
Publicado: (2025)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
por: R, Vinotha, et al.
Publicado: (2024)
por: R, Vinotha, et al.
Publicado: (2024)
Multi-Level Speaker Representation for Target Speaker Extraction
por: Zhang, Ke, et al.
Publicado: (2024)
por: Zhang, Ke, et al.
Publicado: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
por: Yamauchi, Kazuki, et al.
Publicado: (2024)
In This Environment, As That Speaker: A Text-Driven Framework for Multi-Attribute Speech Conversion
por: Jin, Jiawei, et al.
Publicado: (2025)
por: Jin, Jiawei, et al.
Publicado: (2025)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
por: Yong, Zheng-Xin, et al.
Publicado: (2025)
Controllable Accent Normalization via Discrete Diffusion
por: Bai, Qibing, et al.
Publicado: (2026)
por: Bai, Qibing, et al.
Publicado: (2026)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
por: Do, Cong-Thanh, et al.
Publicado: (2024)
por: Do, Cong-Thanh, et al.
Publicado: (2024)
Rethinking Discrete Speech Representation Tokens for Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
Accent-VITS:accent transfer for end-to-end TTS
por: Ma, Linhan, et al.
Publicado: (2023)
por: Ma, Linhan, et al.
Publicado: (2023)
Improving Language Model-Based Zero-Shot Text-to-Speech Synthesis with Multi-Scale Acoustic Prompts
por: Lei, Shun, et al.
Publicado: (2023)
por: Lei, Shun, et al.
Publicado: (2023)
Hierarchical Emotion Prediction and Control in Text-to-Speech Synthesis
por: Inoue, Sho, et al.
Publicado: (2024)
por: Inoue, Sho, et al.
Publicado: (2024)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
por: Kim, Jaeyoung, et al.
Publicado: (2024)
por: Kim, Jaeyoung, et al.
Publicado: (2024)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
por: Chen, Jinming, et al.
Publicado: (2024)
por: Chen, Jinming, et al.
Publicado: (2024)
Boosting Multi-Speaker Expressive Speech Synthesis with Semi-supervised Contrastive Learning
por: Zhu, Xinfa, et al.
Publicado: (2023)
por: Zhu, Xinfa, et al.
Publicado: (2023)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
por: Kang, Jiawen, et al.
Publicado: (2024)
por: Kang, Jiawen, et al.
Publicado: (2024)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
por: Poncelet, Jakob, et al.
Publicado: (2023)
por: Poncelet, Jakob, et al.
Publicado: (2023)
Investigation of Deep Neural Network Acoustic Modelling Approaches for Low Resource Accented Mandarin Speech Recognition
por: Xie, Xurong, et al.
Publicado: (2022)
por: Xie, Xurong, et al.
Publicado: (2022)
Optimizing Multilingual Text-To-Speech with Accents & Emotions
por: Pawar, Pranav, et al.
Publicado: (2025)
por: Pawar, Pranav, et al.
Publicado: (2025)
AVFSNet: Audio-Visual Speech Separation for Flexible Number of Speakers with Multi-Scale and Multi-Task Learning
por: Zhang, Daning, et al.
Publicado: (2025)
por: Zhang, Daning, et al.
Publicado: (2025)
Generating Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems
por: Chen, Zhengyang, et al.
Publicado: (2024)
por: Chen, Zhengyang, et al.
Publicado: (2024)
Prosodically Enhanced Foreign Accent Simulation by Discrete Token-based Resynthesis Only with Native Speech Corpora
por: Onda, Kentaro, et al.
Publicado: (2025)
por: Onda, Kentaro, et al.
Publicado: (2025)
Ejemplares similares
-
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
por: Inoue, Sho, et al.
Publicado: (2024) -
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
por: Melechovsky, Jan, et al.
Publicado: (2024) -
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
por: Mu, Bingshen, et al.
Publicado: (2024) -
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
por: Bai, Qibing, et al.
Publicado: (2026) -
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
por: Cheng, Zhuangfei, et al.
Publicado: (2025)