Salvato in:
| Autore principale: | Yang, Qiaoyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2502.00295 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
di: Zhu, Jian, et al.
Pubblicazione: (2023)
di: Zhu, Jian, et al.
Pubblicazione: (2023)
Open vocabulary keyword spotting through transfer learning from speech synthesis
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
di: V, Kesavaraj, et al.
Pubblicazione: (2024)
A circular microphone array with virtual microphones based on acoustics-informed neural networks
di: Zhao, Sipei, et al.
Pubblicazione: (2024)
di: Zhao, Sipei, et al.
Pubblicazione: (2024)
Boosting keyword spotting through on-device learnable user speech characteristics
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
di: Cioflan, Cristian, et al.
Pubblicazione: (2024)
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)
Effects of automotive microphone frequency response characteristics and noise conditions on speech and ASR quality -- an experimental evaluation
di: Buccoli, Michele, et al.
Pubblicazione: (2025)
di: Buccoli, Michele, et al.
Pubblicazione: (2025)
EBEN: Extreme bandwidth extension network applied to speech signals captured with noise-resilient body-conduction microphones
di: Hauret, Julien, et al.
Pubblicazione: (2022)
di: Hauret, Julien, et al.
Pubblicazione: (2022)
Neural Ambisonics encoding for compact irregular microphone arrays
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
di: Heikkinen, Mikko, et al.
Pubblicazione: (2024)
Binaural rendering from microphone array signals of arbitrary geometry
di: Iijima, Naoto, et al.
Pubblicazione: (2021)
di: Iijima, Naoto, et al.
Pubblicazione: (2021)
Vibration Sensitivity of one-port and two-port MEMS microphones
di: Doyon-D'Amour, Francis, et al.
Pubblicazione: (2024)
di: Doyon-D'Amour, Francis, et al.
Pubblicazione: (2024)
Listening broadband physical model for microphones: a first step
di: Millot, Laurent, et al.
Pubblicazione: (2024)
di: Millot, Laurent, et al.
Pubblicazione: (2024)
Modal smoothing for analysis of room reflections measured with spherical microphone and loudspeaker arrays
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
Design framework for spherical microphone and loudspeaker arrays in a multiple-input multiple-output system
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
di: Morgenstern, Hai, et al.
Pubblicazione: (2024)
Improved direction of arrival estimations with a wearable microphone array for dynamic environments by reliability weighting
di: Mitchell, Daniel A., et al.
Pubblicazione: (2024)
di: Mitchell, Daniel A., et al.
Pubblicazione: (2024)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
Paraformer-v2: An improved non-autoregressive transformer for noise-robust speech recognition
di: An, Keyu, et al.
Pubblicazione: (2024)
di: An, Keyu, et al.
Pubblicazione: (2024)
Interfacing PDM MEMS microphones with PFM spiking systems: Application for Neuromorphic Auditory Sensors
di: Jimenez-Fernandez, Angel, et al.
Pubblicazione: (2019)
di: Jimenez-Fernandez, Angel, et al.
Pubblicazione: (2019)
From the perspective of perceptual speech quality: The robustness of frequency bands to noise
di: Fan, Junyi, et al.
Pubblicazione: (2025)
di: Fan, Junyi, et al.
Pubblicazione: (2025)
A data-driven two-microphone method for in-situ sound absorption measurements
di: Emmerich, Leon, et al.
Pubblicazione: (2025)
di: Emmerich, Leon, et al.
Pubblicazione: (2025)
WavJEPA: Semantic learning unlocks robust audio foundation models for raw waveforms
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
di: Yuksel, Goksenin, et al.
Pubblicazione: (2025)
Hardware-accelerated graph neural networks: an alternative approach for neuromorphic event-based audio classification and keyword spotting on SoC FPGA
di: Jeziorek, Kamil, et al.
Pubblicazione: (2026)
di: Jeziorek, Kamil, et al.
Pubblicazione: (2026)
Improving vision-inspired keyword spotting using dynamic module skipping in streaming conformer encoder
di: Bittar, Alexandre, et al.
Pubblicazione: (2023)
di: Bittar, Alexandre, et al.
Pubblicazione: (2023)
A noise-robust acoustic method for recognizing foraging activities of grazing cattle
di: Martinez-Rau, Luciano S., et al.
Pubblicazione: (2023)
di: Martinez-Rau, Luciano S., et al.
Pubblicazione: (2023)
Cascaded noise reduction and acoustic echo cancellation based on an extended noise reduction
di: Roebben, Arnout, et al.
Pubblicazione: (2024)
di: Roebben, Arnout, et al.
Pubblicazione: (2024)
Towards robust paralinguistic assessment for real-world mobile health (mHealth) monitoring: an initial study of reverberation effects on speech
di: Dineley, Judith, et al.
Pubblicazione: (2023)
di: Dineley, Judith, et al.
Pubblicazione: (2023)
A robust audio deepfake detection system via multi-view feature
di: Yang, Yujie, et al.
Pubblicazione: (2024)
di: Yang, Yujie, et al.
Pubblicazione: (2024)
Discrete Tokens Exhibit Interlanguage Speech Intelligibility Benefit: an Analytical Study Towards Accent-robust ASR Only with Native Speech Data
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
Speech-preserving active noise control: a deep learning approach in reverberant environments
di: Dai, Shuning
Pubblicazione: (2026)
di: Dai, Shuning
Pubblicazione: (2026)
Towards interpretable emotion recognition: Identifying key features with machine learning
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
di: Kaloga, Yacouba, et al.
Pubblicazione: (2025)
A lightweight and robust method for blind wideband-to-fullband extension of speech
di: Büthe, Jan, et al.
Pubblicazione: (2024)
di: Büthe, Jan, et al.
Pubblicazione: (2024)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
di: Yanir, Efrayim, et al.
Pubblicazione: (2025)
Hidden bawls, whispers, and yelps: can text be made to sound more than just its words?
di: Pataca, Caluã de Lacerda, et al.
Pubblicazione: (2022)
di: Pataca, Caluã de Lacerda, et al.
Pubblicazione: (2022)
A microscopic investigation of the effect of random envelope fluctuations on phoneme-in-noise perception
di: Osses, Alejandro, et al.
Pubblicazione: (2024)
di: Osses, Alejandro, et al.
Pubblicazione: (2024)
ComplexDec: A Domain-robust High-fidelity Neural Audio Codec with Complex Spectrum Modeling
di: Wu, Yi-Chiao, et al.
Pubblicazione: (2025)
di: Wu, Yi-Chiao, et al.
Pubblicazione: (2025)
Controllable joint noise reduction and hearing loss compensation using a differentiable auditory model
di: Gonzalez, Philippe, et al.
Pubblicazione: (2025)
di: Gonzalez, Philippe, et al.
Pubblicazione: (2025)
Localizing broadband noise sources using the Loève spectrum and a 2.5D approach
di: Kasess, Christian H., et al.
Pubblicazione: (2026)
di: Kasess, Christian H., et al.
Pubblicazione: (2026)
Towards Decoupling Frontend Enhancement and Backend Recognition in Monaural Robust ASR
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
di: Yang, Yufeng, et al.
Pubblicazione: (2024)
AdaKWS: Towards Robust Keyword Spotting with Test-Time Adaptation
di: Xiao, Yang, et al.
Pubblicazione: (2025)
di: Xiao, Yang, et al.
Pubblicazione: (2025)
Towards Expressive Zero-Shot Speech Synthesis with Hierarchical Prosody Modeling
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
di: Jiang, Yuepeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The taste of IPA: Towards open-vocabulary keyword spotting and forced alignment in any language
di: Zhu, Jian, et al.
Pubblicazione: (2023) -
Open vocabulary keyword spotting through transfer learning from speech synthesis
di: V, Kesavaraj, et al.
Pubblicazione: (2024) -
A circular microphone array with virtual microphones based on acoustics-informed neural networks
di: Zhao, Sipei, et al.
Pubblicazione: (2024) -
Boosting keyword spotting through on-device learnable user speech characteristics
di: Cioflan, Cristian, et al.
Pubblicazione: (2024) -
Towards noise-robust speech inversion through multi-task learning with speech enhancement
di: Tabatabaee, Saba, et al.
Pubblicazione: (2026)