Speech Robust Bench: A Robustness Benchmark For Speech Recognition
Fuente:
arXiv
Salvato in:
| Autori principali: | Shah, Muhammad A., Noguero, David Solans, Heikkila, Mikko A., Raj, Bhiksha, Kourtellis, Nicolas |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Revisiting Acoustic Features for Robust ASR
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
di: Shah, Muhammad A., et al.
Pubblicazione: (2024)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026)
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
di: Waheed, Abdul, et al.
Pubblicazione: (2024)
Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
di: Leem, Seong-Gyun, et al.
Pubblicazione: (2024)
di: Leem, Seong-Gyun, et al.
Pubblicazione: (2024)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)
di: Ngo, Huong, et al.
Pubblicazione: (2025)
Psychoacoustic Challenges Of Speech Enhancement On VoIP Platforms
di: Konan, Joseph, et al.
Pubblicazione: (2023)
di: Konan, Joseph, et al.
Pubblicazione: (2023)
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
WhisperRT -- Turning Whisper into a Causal Streaming Model
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
di: Krichli, Tomer, et al.
Pubblicazione: (2025)
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
RelUNet: Relative Channel Fusion U-Net for Multichannel Speech Enhancement
di: Aldarmaki, Ibrahim, et al.
Pubblicazione: (2024)
di: Aldarmaki, Ibrahim, et al.
Pubblicazione: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
di: Wang, He, et al.
Pubblicazione: (2025)
di: Wang, He, et al.
Pubblicazione: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
di: Rossenbach, Nick, et al.
Pubblicazione: (2024)
Large Language Models are Efficient Learners of Noise-Robust Speech Recognition
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
di: Hu, Yuchen, et al.
Pubblicazione: (2024)
Exploration of Adapter for Noise Robust Automatic Speech Recognition
di: Shi, Hao, et al.
Pubblicazione: (2024)
di: Shi, Hao, et al.
Pubblicazione: (2024)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
di: Jeffries, Nat, et al.
Pubblicazione: (2024)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
di: Chen, Maximillian, et al.
Pubblicazione: (2023)
di: Chen, Maximillian, et al.
Pubblicazione: (2023)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2024)
di: Xu, Hainan, et al.
Pubblicazione: (2024)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
di: Chan, David M., et al.
Pubblicazione: (2024)
di: Chan, David M., et al.
Pubblicazione: (2024)
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
di: Feng, Kexin, et al.
Pubblicazione: (2024)
di: Feng, Kexin, et al.
Pubblicazione: (2024)
Analyzing Speech Unit Selection for Textless Speech-to-Speech Translation
di: Duret, Jarod, et al.
Pubblicazione: (2024)
di: Duret, Jarod, et al.
Pubblicazione: (2024)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
di: Shi, Zhonghao, et al.
Pubblicazione: (2024)
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
di: Choi, Anna Seo Gyeong, et al.
Pubblicazione: (2025)
Pisets: A Robust Speech Recognition System for Lectures and Interviews
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
di: Bondarenko, Ivan, et al.
Pubblicazione: (2026)
TRNet: Two-level Refinement Network leveraging Speech Enhancement for Noise Robust Speech Emotion Recognition
di: Chen, Chengxin, et al.
Pubblicazione: (2024)
di: Chen, Chengxin, et al.
Pubblicazione: (2024)
Less is More: Accurate Speech Recognition & Translation without Web-Scale Data
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
di: Puvvada, Krishna C., et al.
Pubblicazione: (2024)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
di: Hilmes, Benedikt, et al.
Pubblicazione: (2024)
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
di: Wang, Pu, et al.
Pubblicazione: (2026)
di: Wang, Pu, et al.
Pubblicazione: (2026)
Clinical BERTScore: An Improved Measure of Automatic Speech Recognition Performance in Clinical Settings
di: Shor, Joel, et al.
Pubblicazione: (2023)
di: Shor, Joel, et al.
Pubblicazione: (2023)
SimulTron: On-Device Simultaneous Speech to Speech Translation
di: Agranovich, Alex, et al.
Pubblicazione: (2024)
di: Agranovich, Alex, et al.
Pubblicazione: (2024)
Translatotron 3: Speech to Speech Translation with Monolingual Data
di: Nachmani, Eliya, et al.
Pubblicazione: (2023)
di: Nachmani, Eliya, et al.
Pubblicazione: (2023)
Multi-Teacher Language-Aware Knowledge Distillation for Multilingual Speech Emotion Recognition
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
di: Bijoy, Mehedi Hasan, et al.
Pubblicazione: (2025)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
di: Mi, Jinyi, et al.
Pubblicazione: (2024)
di: Mi, Jinyi, et al.
Pubblicazione: (2024)
BanglaRobustNet: A Hybrid Denoising-Attention Architecture for Robust Bangla Speech Recognition
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2026)
di: Ridoy, Md Sazzadul Islam, et al.
Pubblicazione: (2026)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
Automatic Speech Recognition for Hindi
di: Saha, Anish, et al.
Pubblicazione: (2024)
di: Saha, Anish, et al.
Pubblicazione: (2024)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Revisiting Acoustic Features for Robust ASR
di: Shah, Muhammad A., et al.
Pubblicazione: (2024) -
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
di: Wang, Chien-Chun, et al.
Pubblicazione: (2026) -
What Do Speech Foundation Models Not Learn About Speech?
di: Waheed, Abdul, et al.
Pubblicazione: (2024) -
Describe Where You Are: Improving Noise-Robustness for Speech Emotion Recognition with Text Description of the Environment
di: Leem, Seong-Gyun, et al.
Pubblicazione: (2024) -
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
di: Ngo, Huong, et al.
Pubblicazione: (2025)