Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
Fuente:
arXiv
Salvato in:
| Autori principali: | Do, Cong-Thanh, Imai, Shuhei, Doddipatla, Rama, Hain, Thomas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024)
di: Li, Mohan, et al.
Pubblicazione: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
di: Nespoli, Francesco, et al.
Pubblicazione: (2024)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
di: Park, Chanho, et al.
Pubblicazione: (2024)
di: Park, Chanho, et al.
Pubblicazione: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2025)
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2025)
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
di: Sasu, David, et al.
Pubblicazione: (2025)
di: Sasu, David, et al.
Pubblicazione: (2025)
MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent Conversion
di: Inoue, Sho, et al.
Pubblicazione: (2024)
di: Inoue, Sho, et al.
Pubblicazione: (2024)
Training Data Augmentation for Dysarthric Automatic Speech Recognition by Text-to-Dysarthric-Speech Synthesis
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
di: Leung, Wing-Zin, et al.
Pubblicazione: (2024)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023)
di: Park, Chanho, et al.
Pubblicazione: (2023)
Multi-Scale Accent Modeling and Disentangling for Multi-Speaker Multi-Accent Text-to-Speech Synthesis
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
di: Zhou, Xuehao, et al.
Pubblicazione: (2024)
Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
di: Sameti, Mohammad Hossein, et al.
Pubblicazione: (2025)
di: Sameti, Mohammad Hossein, et al.
Pubblicazione: (2025)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
di: Dossou, Bonaventure F. P.
Pubblicazione: (2023)
di: Dossou, Bonaventure F. P.
Pubblicazione: (2023)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
di: Bafna, Niyati, et al.
Pubblicazione: (2025)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
UtterTune: LoRA-Based Target-Language Pronunciation Edit and Control in Multilingual Text-to-Speech
di: Kato, Shuhei
Pubblicazione: (2025)
di: Kato, Shuhei
Pubblicazione: (2025)
VoxEmo: Benchmarking Speech Emotion Recognition with Speech LLMs
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
di: Zhang, Hezhao, et al.
Pubblicazione: (2026)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
Rethinking Discrete Speech Representation Tokens for Accent Generation
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2026)
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2026)
Performant ASR Models for Medical Entities in Accented Speech
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
di: Afonja, Tejumade, et al.
Pubblicazione: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
di: Shivakumar, Prashanth Gurunath, et al.
Pubblicazione: (2024)
Improving Speech Emotion Recognition in Under-Resourced Languages via Speech-to-Speech Translation with Bootstrapping Data Selection
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
di: Lin, Hsi-Che, et al.
Pubblicazione: (2024)
Towards Unsupervised Speech Recognition Without Pronunciation Models
di: Ni, Junrui, et al.
Pubblicazione: (2024)
di: Ni, Junrui, et al.
Pubblicazione: (2024)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
Empowering Communication: Speech Technology for Indian and Western Accents through AI-powered Speech Synthesis
di: R, Vinotha, et al.
Pubblicazione: (2024)
di: R, Vinotha, et al.
Pubblicazione: (2024)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
di: Guo, Jiaxin, et al.
Pubblicazione: (2024)
Hallucination in Perceptual Metric-Driven Speech Enhancement Networks
di: Close, George, et al.
Pubblicazione: (2024)
di: Close, George, et al.
Pubblicazione: (2024)
Accented Text-to-Speech Synthesis with a Conditional Variational Autoencoder
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
di: Melechovsky, Jan, et al.
Pubblicazione: (2022)
Clustering and Mining Accented Speech for Inclusive and Fair Speech Recognition
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2024)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
di: Shen, Peng, et al.
Pubblicazione: (2025)
di: Shen, Peng, et al.
Pubblicazione: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Improving Speech-based Emotion Recognition with Contextual Utterance Analysis and LLMs
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
di: Zhang, Enshi, et al.
Pubblicazione: (2024)
Using Speech Foundational Models in Loss Functions for Hearing Aid Speech Enhancement
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
di: Sutherland, Robert, et al.
Pubblicazione: (2024)
Improving Whisper's Recognition Performance for Under-Represented Language Kazakh Leveraging Unpaired Speech and Text
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
di: Li, Jinpeng, et al.
Pubblicazione: (2024)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
Speaker-Aware Simulation Improves Conversational Speech Recognition
di: Gedeon, Máté, et al.
Pubblicazione: (2026)
di: Gedeon, Máté, et al.
Pubblicazione: (2026)
Continuous Speech Tokenizer in Text To Speech
di: Li, Yixing, et al.
Pubblicazione: (2024)
di: Li, Yixing, et al.
Pubblicazione: (2024)
Adaptive Inner Speech-Text Alignment for LLM-based Speech Translation
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
di: Liu, Henglyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
WHISMA: A Speech-LLM to Perform Zero-shot Spoken Language Understanding
di: Li, Mohan, et al.
Pubblicazione: (2024) -
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
di: Nespoli, Francesco, et al.
Pubblicazione: (2024) -
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
di: Meghanani, Amit, et al.
Pubblicazione: (2024) -
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024) -
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026)