Few-Shot Accent Synthesis for ASR with LLM-Guided Phoneme Editing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Halychanskyi, Yurii, Bozdag, Nimet Beyza, Hasegawa-Johnson, Mark, Hakkani-Tür, Dilek, Kindratenko, Volodymyr |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2025)
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2025)
Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2026)
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2026)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
Language Specific Knowledge: Do Models Know Better in X than in English?
von: Agarwal, Ishika, et al.
Veröffentlicht: (2025)
von: Agarwal, Ishika, et al.
Veröffentlicht: (2025)
DialDefer: A Framework for Detecting and Mitigating LLM Dialogic Deference
von: Rabbani, Parisa, et al.
Veröffentlicht: (2026)
von: Rabbani, Parisa, et al.
Veröffentlicht: (2026)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
von: Owodunni, Abraham Toluwase, et al.
Veröffentlicht: (2024)
von: Owodunni, Abraham Toluwase, et al.
Veröffentlicht: (2024)
Effects of Speaker Count, Duration, and Accent Diversity on Zero-Shot Accent Robustness in Low-Resource ASR
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
von: Yong, Zheng-Xin, et al.
Veröffentlicht: (2025)
Contrastive Regularization for Accent-Robust ASR
von: Thai, Van-Phat, et al.
Veröffentlicht: (2026)
von: Thai, Van-Phat, et al.
Veröffentlicht: (2026)
Towards Robust Dysarthric Speech Recognition: LLM-Agent Post-ASR Correction Beyond WER
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2026)
Multimodal Consistency-Guided Reference-Free Data Selection for ASR Accent Adaptation
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
von: Lei, Ligong, et al.
Veröffentlicht: (2026)
AURA: A Diagnostic Framework for Tracking User Satisfaction of Interactive Planning Agents
von: Kim, Takyoung, et al.
Veröffentlicht: (2025)
von: Kim, Takyoung, et al.
Veröffentlicht: (2025)
Performant ASR Models for Medical Entities in Accented Speech
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
von: Afonja, Tejumade, et al.
Veröffentlicht: (2024)
End-to-End Zero-Shot Voice Conversion with Location-Variable Convolutions
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
von: Kang, Wonjune, et al.
Veröffentlicht: (2022)
Too Polite to Disagree: Understanding Sycophancy Propagation in Multi-Agent Systems
von: Kasprova, Vira, et al.
Veröffentlicht: (2026)
von: Kasprova, Vira, et al.
Veröffentlicht: (2026)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2024)
Must Read: A Comprehensive Survey of Computational Persuasion
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025)
DITTO: Data-efficient and Fair Targeted Subset Selection for ASR Accent Adaptation
von: Kothawade, Suraj, et al.
Veröffentlicht: (2021)
von: Kothawade, Suraj, et al.
Veröffentlicht: (2021)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
von: Inoue, Sho, et al.
Veröffentlicht: (2024)
Data-Efficient ASR Personalization for Non-Normative Speech Using an Uncertainty-Based Phoneme Difficulty Score for Guided Sampling
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
von: Pokel, Niclas, et al.
Veröffentlicht: (2025)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
von: Zhou, Xuehao, et al.
Veröffentlicht: (2024)
PARCO: Phoneme-Augmented Robust Contextual ASR via Contrastive Entity Disambiguation
von: He, Jiajun, et al.
Veröffentlicht: (2025)
von: He, Jiajun, et al.
Veröffentlicht: (2025)
ProGRes: Prompted Generative Rescoring on ASR n-Best
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
von: Tur, Ada Defne, et al.
Veröffentlicht: (2024)
CosyAccent: Duration-Controllable Accent Normalization Using Source-Synthesis Training Data
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
von: Bai, Qibing, et al.
Veröffentlicht: (2026)
SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
von: Cheng, Zhuangfei, et al.
Veröffentlicht: (2025)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
von: Zhong, Jinzuomu, et al.
Veröffentlicht: (2025)
MetaSICL: Adapting Audiroty LLM via Meta Speech In-Context Learning
von: Zheng, Haolong, et al.
Veröffentlicht: (2026)
von: Zheng, Haolong, et al.
Veröffentlicht: (2026)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
von: Dossou, Bonaventure F. P.
Veröffentlicht: (2023)
In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
von: Fan, Xulin, et al.
Veröffentlicht: (2026)
Multi-Accent Mandarin Dry-Vocal Singing Dataset: Benchmark for Singing Accent Recognition
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
von: Wang, Zihao, et al.
Veröffentlicht: (2025)
LLaDA-TTS: Unifying Speech Synthesis and Zero-Shot Editing via Masked Diffusion Modeling
von: Fan, Xiaoyu, et al.
Veröffentlicht: (2026)
von: Fan, Xiaoyu, et al.
Veröffentlicht: (2026)
Advanced Modeling of Interlanguage Speech Intelligibility Benefit with L1-L2 Multi-Task Learning Using Differentiable K-Means for Accent-Robust Discrete Token-Based ASR
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
von: Onda, Kentaro, et al.
Veröffentlicht: (2026)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
von: Onda, Kentaro, et al.
Veröffentlicht: (2025)
Few-Shot Keyword Spotting from Mixed Speech
von: Yuan, Junming, et al.
Veröffentlicht: (2024)
von: Yuan, Junming, et al.
Veröffentlicht: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
von: Nespoli, Francesco, et al.
Veröffentlicht: (2024)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
von: Zheng, Xiuwen, et al.
Veröffentlicht: (2024)
Enhancing Child Vocalization Classification with Phonetically-Tuned Embeddings for Assisting Autism Diagnosis
von: Li, Jialu, et al.
Veröffentlicht: (2023)
von: Li, Jialu, et al.
Veröffentlicht: (2023)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
von: Melechovsky, Jan, et al.
Veröffentlicht: (2022)
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages
von: Maltais, Marie, et al.
Veröffentlicht: (2026)
von: Maltais, Marie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FAC-FACodec: Controllable Zero-Shot Foreign Accent Conversion with Factorized Speech Codec
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2025) -
Accent Conversion: A Problem-Driven Survey of Sociolinguistic and Technical Constraints
von: Halychanskyi, Yurii, et al.
Veröffentlicht: (2026) -
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
von: Rabbani, Parisa, et al.
Veröffentlicht: (2025) -
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
von: Bozdag, Nimet Beyza, et al.
Veröffentlicht: (2025) -
Language Specific Knowledge: Do Models Know Better in X than in English?
von: Agarwal, Ishika, et al.
Veröffentlicht: (2025)