AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Jinzuomu, Richmond, Korin, Su, Zhiba, Sun, Siqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
von: Owodunni, Abraham Toluwase, et al.
Veröffentlicht: (2024)
von: Owodunni, Abraham Toluwase, et al.
Veröffentlicht: (2024)
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
von: Sun, Siqi, et al.
Veröffentlicht: (2024)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2023)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
von: Bafna, Niyati, et al.
Veröffentlicht: (2025)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
von: Sasu, David, et al.
Veröffentlicht: (2025)
von: Sasu, David, et al.
Veröffentlicht: (2025)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuan-Nam, et al.
Veröffentlicht: (2025)
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
von: Sameti, Mohammad Hossein, et al.
Veröffentlicht: (2025)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
Segmentation-Variant Codebooks for Preservation of Paralinguistic and Prosodic Information
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
von: Sanders, Nicholas, et al.
Veröffentlicht: (2025)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
On the Relationship between Accent Strength and Articulatory Features
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
von: Huang, Kevin, et al.
Veröffentlicht: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
von: Nguyen, Tuan Nam, et al.
Veröffentlicht: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
von: Huang, Wen-Chin, et al.
Veröffentlicht: (2026)
Exploring Acoustic Similarity in Emotional Speech and Music via Self-Supervised Representations
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
von: Sun, Yujia, et al.
Veröffentlicht: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
Multilingual and Multi-Accent Jailbreaking of Audio LLMs
von: Roh, Jaechul, et al.
Veröffentlicht: (2025)
von: Roh, Jaechul, et al.
Veröffentlicht: (2025)
Accent-VITS:accent transfer for end-to-end TTS
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
von: Ma, Linhan, et al.
Veröffentlicht: (2023)
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
Convert and Speak: Zero-shot Accent Conversion with Minimum Supervision
von: Jia, Zhijun, et al.
Veröffentlicht: (2024)
von: Jia, Zhijun, et al.
Veröffentlicht: (2024)
Study on the Fairness of Speaker Verification Systems on Underrepresented Accents in English
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
von: Estevez, Mariel, et al.
Veröffentlicht: (2022)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
von: Doan, Khai Duy, et al.
Veröffentlicht: (2024)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation
von: Kothawade, Suraj, et al.
Veröffentlicht: (2021)
von: Kothawade, Suraj, et al.
Veröffentlicht: (2021)
Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
von: Bai, Qibing, et al.
Veröffentlicht: (2025)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
von: Mu, Bingshen, et al.
Veröffentlicht: (2024)
GSA-TTS : Toward Zero-Shot Speech Synthesis based on Gradual Style Adaptor
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
von: Lee, Seokgi, et al.
Veröffentlicht: (2025)
Zero-Shot Text-to-Speech for Vietnamese
von: Vu, Thi, et al.
Veröffentlicht: (2025)
von: Vu, Thi, et al.
Veröffentlicht: (2025)
Controllable Accent Normalization via Discrete Diffusion
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
Not that Groove: Zero-Shot Symbolic Music Editing
von: Zhang, Li
Veröffentlicht: (2025)
von: Zhang, Li
Veröffentlicht: (2025)
Ähnliche Einträge
-
Rethinking Discrete Speech Representation Tokens for Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2026) -
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025) -
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
von: Owodunni, Abraham Toluwase, et al.
Veröffentlicht: (2024) -
Acquiring Pronunciation Knowledge from Transcribed Speech Audio via Multi-task Learning
von: Sun, Siqi, et al.
Veröffentlicht: (2024) -
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)