Accent-Invariant Automatic Speech Recognition via Saliency-Driven Spectrogram Masking
Fuente:
arXiv
Guardado en:
| Autores principales: | Sameti, Mohammad Hossein, Moridani, Sepehr Harfi, Zarean, Ali, Sameti, Hossein |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
por: Sasu, David, et al.
Publicado: (2025)
por: Sasu, David, et al.
Publicado: (2025)
Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition
por: Iakovenko, Olga, et al.
Publicado: (2024)
por: Iakovenko, Olga, et al.
Publicado: (2024)
Where Are You From? Let Me Guess! Subdialect Recognition of Speeches in Sorani Kurdish
por: Isam, Sana, et al.
Publicado: (2024)
por: Isam, Sana, et al.
Publicado: (2024)
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
por: Dossou, Bonaventure F. P.
Publicado: (2023)
por: Dossou, Bonaventure F. P.
Publicado: (2023)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
por: Do, Cong-Thanh, et al.
Publicado: (2024)
por: Do, Cong-Thanh, et al.
Publicado: (2024)
LID Models are Actually Accent Classifiers: Implications and Solutions for LID on Accented Speech
por: Bafna, Niyati, et al.
Publicado: (2025)
por: Bafna, Niyati, et al.
Publicado: (2025)
Automatic Speech Recognition for Hindi
por: Saha, Anish, et al.
Publicado: (2024)
por: Saha, Anish, et al.
Publicado: (2024)
Pairwise Evaluation of Accent Similarity in Speech Synthesis
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
por: Zhong, Jinzuomu, et al.
Publicado: (2025)
Rethinking Discrete Speech Representation Tokens for Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
por: Zhong, Jinzuomu, et al.
Publicado: (2026)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
SPES: Spectrogram Perturbation for Explainable Speech-to-Text Generation
por: Fucci, Dennis, et al.
Publicado: (2024)
por: Fucci, Dennis, et al.
Publicado: (2024)
Dynamic Data Pruning for Automatic Speech Recognition
por: Xiao, Qiao, et al.
Publicado: (2024)
por: Xiao, Qiao, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
por: Hu, Jiliang, et al.
Publicado: (2025)
por: Hu, Jiliang, et al.
Publicado: (2025)
Benchmarking Automatic Speech Recognition Models for African Languages
por: Nahabwe, Alvin, et al.
Publicado: (2025)
por: Nahabwe, Alvin, et al.
Publicado: (2025)
Exploring Gender Disparities in Automatic Speech Recognition Technology
por: ElGhazaly, Hend, et al.
Publicado: (2025)
por: ElGhazaly, Hend, et al.
Publicado: (2025)
Exploration of Adapter for Noise Robust Automatic Speech Recognition
por: Shi, Hao, et al.
Publicado: (2024)
por: Shi, Hao, et al.
Publicado: (2024)
Automatic Speech Recognition for Biomedical Data in Bengali Language
por: Kabir, Shariar, et al.
Publicado: (2024)
por: Kabir, Shariar, et al.
Publicado: (2024)
Zero Shot Text to Speech Augmentation for Automatic Speech Recognition on Low-Resource Accented Speech Corpora
por: Nespoli, Francesco, et al.
Publicado: (2024)
por: Nespoli, Francesco, et al.
Publicado: (2024)
Persian Musical Instruments Classification Using Polyphonic Data Augmentation
por: Esfangereh, Diba Hadi, et al.
Publicado: (2025)
por: Esfangereh, Diba Hadi, et al.
Publicado: (2025)
Persian MusicGen: A Large-Scale Dataset and Culturally-Aware Generative Model for Persian Music
por: Sameti, Mohammad Hossein, et al.
Publicado: (2026)
por: Sameti, Mohammad Hossein, et al.
Publicado: (2026)
AccentFold: A Journey through African Accents for Zero-Shot ASR Adaptation to Target Accents
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
por: Owodunni, Abraham Toluwase, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
Supporting SENCOTEN Language Documentation Efforts with Automatic Speech Recognition
por: Geng, Mengzhe, et al.
Publicado: (2025)
por: Geng, Mengzhe, et al.
Publicado: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
por: Hu, Ke, et al.
Publicado: (2025)
por: Hu, Ke, et al.
Publicado: (2025)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary Prediction and Activation
por: Lin, Zhennan, et al.
Publicado: (2025)
por: Lin, Zhennan, et al.
Publicado: (2025)
Evaluating Automatic Speech Recognition Systems for Korean Meteorological Experts
por: Park, ChaeHun, et al.
Publicado: (2024)
por: Park, ChaeHun, et al.
Publicado: (2024)
Transliterated Zero-Shot Domain Adaptation for Automatic Speech Recognition
por: Zhu, Han, et al.
Publicado: (2024)
por: Zhu, Han, et al.
Publicado: (2024)
UCorrect: An Unsupervised Framework for Automatic Speech Recognition Error Correction
por: Guo, Jiaxin, et al.
Publicado: (2024)
por: Guo, Jiaxin, et al.
Publicado: (2024)
Enhancing Indonesian Automatic Speech Recognition: Evaluating Multilingual Models with Diverse Speech Variabilities
por: Adila, Aulia, et al.
Publicado: (2024)
por: Adila, Aulia, et al.
Publicado: (2024)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
por: Hsu, Ming-Hao, et al.
Publicado: (2024)
por: Hsu, Ming-Hao, et al.
Publicado: (2024)
AccentBox: Towards High-Fidelity Zero-Shot Accent Generation
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
por: Zhong, Jinzuomu, et al.
Publicado: (2024)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
por: Venkateswaran, Nitin, et al.
Publicado: (2025)
Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language
por: Abu, Turi, et al.
Publicado: (2025)
por: Abu, Turi, et al.
Publicado: (2025)
Automatic Speech Recognition for Non-Native English: Accuracy and Disfluency Handling
por: McGuire, Michael
Publicado: (2025)
por: McGuire, Michael
Publicado: (2025)
A Deep Learning Automatic Speech Recognition Model for Shona Language
por: Sirora, Leslie Wellington, et al.
Publicado: (2025)
por: Sirora, Leslie Wellington, et al.
Publicado: (2025)
Automatic Speech Recognition System-Independent Word Error Rate Estimation
por: Park, Chanho, et al.
Publicado: (2024)
por: Park, Chanho, et al.
Publicado: (2024)
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
por: Frieske, Rita, et al.
Publicado: (2024)
por: Frieske, Rita, et al.
Publicado: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
Ejemplares similares
-
Pitch Accent Detection improves Pretrained Automatic Speech Recognition
por: Sasu, David, et al.
Publicado: (2025) -
Convolutional Variational Autoencoders for Spectrogram Compression in Automatic Speech Recognition
por: Iakovenko, Olga, et al.
Publicado: (2024) -
Where Are You From? Let Me Guess! Subdialect Recognition of Speeches in Sorani Kurdish
por: Isam, Sana, et al.
Publicado: (2024) -
Advancing African-Accented Speech Recognition: Epistemic Uncertainty-Driven Data Selection for Generalizable ASR Models
por: Dossou, Bonaventure F. P.
Publicado: (2023) -
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
por: Do, Cong-Thanh, et al.
Publicado: (2024)