Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Buccoli, Michele, Du, Yu, Soendergaard, Jacob, Cazzaniga, Simone Shawn |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
di: Hauret, Julien, et al.
Pubblicazione: (2022)
di: Hauret, Julien, et al.
Pubblicazione: (2022)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
di: Fan, Junyi, et al.
Pubblicazione: (2025)
di: Fan, Junyi, et al.
Pubblicazione: (2025)
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
di: Yang, Qiaoyu
Pubblicazione: (2025)
di: Yang, Qiaoyu
Pubblicazione: (2025)
A circular microphone array with virtual microphones based on acoustics-informed neural networks
di: Zhao, Sipei, et al.
Pubblicazione: (2024)
di: Zhao, Sipei, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Neural Ambisonics encoding for compact irregular microphone arrays
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
Binaural rendering from microphone array signals of arbitrary geometry
di: Iijima, Naoto, et al.
Pubblicazione: (2021)
di: Iijima, Naoto, et al.
Pubblicazione: (2021)
Vibration Sensitivity of one-port and two-port MEMS microphones
di: Doyon-D'Amour, Francis, et al.
Pubblicazione: (2024)
di: Doyon-D'Amour, Francis, et al.
Pubblicazione: (2024)
Listening broadband physical model for microphones: a first step
di: Millot, Laurent, et al.
Pubblicazione: (2024)
di: Millot, Laurent, et al.
Pubblicazione: (2024)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
Modal smoothing for analysis of room reflections measured with spherical microphone and loudspeaker arrays
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Design framework for spherical microphone and loudspeaker arrays in a multiple-input multiple-output system
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
di: Mitchell, Daniel A., et al.
Pubblicazione: (2024)
di: Mitchell, Daniel A., et al.
Pubblicazione: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
di: Watanabe, Aya, et al.
Pubblicazione: (2024)
Determined blind source separation via modeling adjacent frequency band correlations in speech signals
di: Wang, Jianyu, et al.
Pubblicazione: (2025)
di: Wang, Jianyu, et al.
Pubblicazione: (2025)
Using RLHF to align speech enhancement approaches to mean-opinion quality scores
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
di: Kumar, Anurag, et al.
Pubblicazione: (2024)
Ensemble of classifiers for speech evaluation
di: Belokrylov, G., et al.
Pubblicazione: (2024)
di: Belokrylov, G., et al.
Pubblicazione: (2024)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
di: Raissi, Tina, et al.
Pubblicazione: (2024)
di: Raissi, Tina, et al.
Pubblicazione: (2024)
Interfacing PDM MEMS microphones with PFM spiking systems: Application for Neuromorphic Auditory Sensors
di: Jimenez-Fernandez, Angel, et al.
Pubblicazione: (2019)
di: Jimenez-Fernandez, Angel, et al.
Pubblicazione: (2019)
Target Speaker ASR with Whisper
di: Polok, Alexander, et al.
Pubblicazione: (2024)
di: Polok, Alexander, et al.
Pubblicazione: (2024)
Index-ASR Technical Report
di: Song, Zheshu, et al.
Pubblicazione: (2025)
di: Song, Zheshu, et al.
Pubblicazione: (2025)
On the relationship between speech and hearing
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
di: Umesh, Srinivasan, et al.
Pubblicazione: (2024)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Efficient Scaling for LLM-based ASR
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
di: Mu, Bingshen, et al.
Pubblicazione: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
di: Ducorroy, Alexandre, et al.
Pubblicazione: (2025)
Detecting gamma-band responses to the speech envelope for the ICASSP 2024 Auditory EEG Decoding Signal Processing Grand Challenge
di: Thornton, Mike, et al.
Pubblicazione: (2024)
di: Thornton, Mike, et al.
Pubblicazione: (2024)
Overlap-Adaptive Hybrid Speaker Diarization and ASR-Aware Observation Addition for MISP 2025 Challenge
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
di: Huang, Shangkun, et al.
Pubblicazione: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
di: Kesiraju, Santosh, et al.
Pubblicazione: (2023)
The USTC-NERCSLIP Systems for The ICMC-ASR Challenge
di: Wu, Minghui, et al.
Pubblicazione: (2024)
di: Wu, Minghui, et al.
Pubblicazione: (2024)
An interpretable speech foundation model for depression detection by revealing prediction-relevant acoustic features from long speech
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
di: Deng, Qingkun, et al.
Pubblicazione: (2024)
LUPET: Incorporating Hierarchical Information Path into Multilingual ASR
di: Liu, Wei, et al.
Pubblicazione: (2024)
di: Liu, Wei, et al.
Pubblicazione: (2024)
persoDA: Personalized Data Augmentation for Personalized ASR
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
di: Parada, Pablo Peso, et al.
Pubblicazione: (2025)
Speaker Adaptation for Quantised End-to-End ASR Models
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
di: Zhao, Qiuming, et al.
Pubblicazione: (2024)
Comparative Analysis of ASR Methods for Speech Deepfake Detection
di: Salvi, Davide, et al.
Pubblicazione: (2024)
di: Salvi, Davide, et al.
Pubblicazione: (2024)
Consistency Based Unsupervised Self-training For ASR Personalisation
di: Zhang, Jisi, et al.
Pubblicazione: (2024)
di: Zhang, Jisi, et al.
Pubblicazione: (2024)
Joint ASR and Speaker Role Tagging with Serialized Output Training
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
Learning When to Trust Which Teacher for Weakly Supervised ASR
di: Agrawal, Aakriti, et al.
Pubblicazione: (2023)
di: Agrawal, Aakriti, et al.
Pubblicazione: (2023)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
di: He, Xiluo, et al.
Pubblicazione: (2025)
di: He, Xiluo, et al.
Pubblicazione: (2025)
Technical Report: A Practical Guide to Kaldi ASR Optimization
di: Hong, Mengze, et al.
Pubblicazione: (2025)
di: Hong, Mengze, et al.
Pubblicazione: (2025)
Documenti analoghi
-
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
di: Hauret, Julien, et al.
Pubblicazione: (2022) -
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
di: Fan, Junyi, et al.
Pubblicazione: (2025) -
Toward noise-robust whisper keyword spotting on headphones with in-earcup microphone and curriculum learning
di: Yang, Qiaoyu
Pubblicazione: (2025) -
A circular microphone array with virtual microphones based on acoustics-informed neural networks
di: Zhao, Sipei, et al.
Pubblicazione: (2024) -
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)