Salvato in:
| Autori principali: | Yang, Mu, Hansen, John H. L. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2603.05977 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Activation Steering for Accent Adaptation in Speech Foundation Models
di: Sun, Jinuo, et al.
Pubblicazione: (2026)
di: Sun, Jinuo, et al.
Pubblicazione: (2026)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
di: Wang, Wenbin, et al.
Pubblicazione: (2024)
di: Wang, Wenbin, et al.
Pubblicazione: (2024)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2024)
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2024)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
Debatts: Zero-Shot Debating Text-to-Speech Synthesis
di: Huang, Yiqiao, et al.
Pubblicazione: (2024)
di: Huang, Yiqiao, et al.
Pubblicazione: (2024)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
di: Owodunni, Abraham Toluwase, et al.
Pubblicazione: (2024)
di: Owodunni, Abraham Toluwase, et al.
Pubblicazione: (2024)
Incremental Disentanglement for Environment-Aware Zero-Shot Text-to-Speech Synthesis
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
di: Lu, Ye-Xin, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech from Continuous Text Streams
di: Dang, Trung, et al.
Pubblicazione: (2024)
di: Dang, Trung, et al.
Pubblicazione: (2024)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
di: Cheng, Zhuangfei, et al.
Pubblicazione: (2025)
di: Cheng, Zhuangfei, et al.
Pubblicazione: (2025)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
di: Zhang, Bowen, et al.
Pubblicazione: (2025)
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
di: Yang, Mu, et al.
Pubblicazione: (2024)
di: Yang, Mu, et al.
Pubblicazione: (2024)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
di: Yong, Zheng-Xin, et al.
Pubblicazione: (2025)
Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition
di: Yang, Mu, et al.
Pubblicazione: (2025)
di: Yang, Mu, et al.
Pubblicazione: (2025)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
di: R, Vinotha, et al.
Pubblicazione: (2024)
di: R, Vinotha, et al.
Pubblicazione: (2024)
Zero-Shot Text-to-Speech for Vietnamese
di: Vu, Thi, et al.
Pubblicazione: (2025)
di: Vu, Thi, et al.
Pubblicazione: (2025)
Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
di: Wang, Tianrui, et al.
Pubblicazione: (2025)
Advanced Zero-Shot Text-to-Speech for Background Removal and Preservation with Controllable Masked Speech Prediction
di: Zhang, Leying, et al.
Pubblicazione: (2025)
di: Zhang, Leying, et al.
Pubblicazione: (2025)
MobileSpeech: A Fast and High-Fidelity Framework for Mobile Zero-Shot Text-to-Speech
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
di: Ji, Shengpeng, et al.
Pubblicazione: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
DART: Disentanglement of Accent and Speaker Representation in Multispeaker Text-to-Speech
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
Adversarial Attacks and Robust Defenses in Speaker Embedding based Zero-Shot Text-to-Speech System
di: Li, Ze, et al.
Pubblicazione: (2024)
di: Li, Ze, et al.
Pubblicazione: (2024)
Mixture of LoRA Experts with Multi-Modal and Multi-Granularity LLM Generative Error Correction for Accented Speech Recognition
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
di: Xie, Tianxin, et al.
Pubblicazione: (2025)
Advancing Zero-Shot Open-Set Speech Deepfake Source Tracing
di: Chhibber, Manasi, et al.
Pubblicazione: (2025)
di: Chhibber, Manasi, et al.
Pubblicazione: (2025)
DAT-CFTNet: Speech Enhancement for Cochlear Implant Recipients using Attention-based Dual-Path Recurrent Neural Network
di: Mamun, Nursadul, et al.
Pubblicazione: (2026)
di: Mamun, Nursadul, et al.
Pubblicazione: (2026)
MMGER: Multi-modal and Multi-granularity Generative Error Correction with LLM for Joint Accent and Speech Recognition
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
di: Mu, Bingshen, et al.
Pubblicazione: (2024)
Analyzing the Impact of Accent on English Speech: Acoustic and Articulatory Perspectives
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
di: Premananth, Gowtham, et al.
Pubblicazione: (2025)
Towards Zero-Shot Text-To-Speech for Arabic Dialects
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
di: Doan, Khai Duy, et al.
Pubblicazione: (2024)
ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching
di: Zhu, Han, et al.
Pubblicazione: (2025)
di: Zhu, Han, et al.
Pubblicazione: (2025)
CosyEdit: Unlocking End-to-End Speech Editing Capability from Zero-Shot Text-to-Speech Models
di: Chen, Junyang, et al.
Pubblicazione: (2026)
di: Chen, Junyang, et al.
Pubblicazione: (2026)
CodecMOS-Accent: A MOS Benchmark of Resynthesized and TTS Speech from Neural Codecs Across English Accents
di: Huang, Wen-Chin, et al.
Pubblicazione: (2026)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2026)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
di: Do, Cong-Thanh, et al.
Pubblicazione: (2024)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
di: Kim, Jaehyeon, et al.
Pubblicazione: (2024)
Parallel GPT: Harmonizing the Independence and Interdependence of Acoustic and Semantic Information for Zero-Shot Text-to-Speech
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
di: Xing, Jingyuan, et al.
Pubblicazione: (2025)
Accent Conversion in Text-To-Speech Using Multi-Level VAE and Adversarial Training
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
di: Melechovsky, Jan, et al.
Pubblicazione: (2024)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
Emotion-Aware Prefix: Towards Explicit Emotion Control in Voice Conversion Models
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
di: Yang, Haoyuan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Activation Steering for Accent Adaptation in Speech Foundation Models
di: Sun, Jinuo, et al.
Pubblicazione: (2026) -
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024) -
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
di: Inoue, Sho, et al.
Pubblicazione: (2024) -
GLOBE: A High-quality English Corpus with Global Accents for Zero-shot Speaker Adaptive Text-to-Speech
di: Wang, Wenbin, et al.
Pubblicazione: (2024) -
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2024)