Loudspeaker Beamforming to Enhance Speech Recognition Performance of Voice Driven Applications
Fuente:
arXiv
Salvato in:
| Autori principali: | de Groot, Dimme, Karslioglu, Baturalp, Scharenborg, Odette, Martinez, Jorge |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
di: de Groot, Dimme, et al.
Pubblicazione: (2025)
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
di: de Groot, Dimme, et al.
Pubblicazione: (2026)
di: de Groot, Dimme, et al.
Pubblicazione: (2026)
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023)
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
di: Patel, Tanvina, et al.
Pubblicazione: (2025)
di: Patel, Tanvina, et al.
Pubblicazione: (2025)
Constant Directivity Loudspeaker Beamforming
di: Luo, Yuancheng
Pubblicazione: (2024)
di: Luo, Yuancheng
Pubblicazione: (2024)
Performance of Objective Speech Quality Metrics on Languages Beyond Validation Data: A Study of Turkish and Korean
di: Perez, Javier, et al.
Pubblicazione: (2025)
di: Perez, Javier, et al.
Pubblicazione: (2025)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
The Multimodal Information Based Speech Processing (MISP) 2025 Challenge: Audio-Visual Diarization and Recognition
di: Gao, Ming, et al.
Pubblicazione: (2025)
di: Gao, Ming, et al.
Pubblicazione: (2025)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
Statistical Beamformer Exploiting Non-stationarity and Sparsity with Spatially Constrained ICA for Robust Speech Recognition
di: Shin, Ui-Hyeop, et al.
Pubblicazione: (2023)
di: Shin, Ui-Hyeop, et al.
Pubblicazione: (2023)
VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via Autoguidance
di: Yeom, Jiheum, et al.
Pubblicazione: (2024)
di: Yeom, Jiheum, et al.
Pubblicazione: (2024)
Auden-Voice: General-Purpose Voice Encoder for Speech and Language Understanding
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
di: Huo, Mingyue, et al.
Pubblicazione: (2025)
Attention-Based Beamformer For Multi-Channel Speech Enhancement
di: Bai, Jinglin, et al.
Pubblicazione: (2024)
di: Bai, Jinglin, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
End-to-End Integration of Speech Emotion Recognition with Voice Activity Detection using Self-Supervised Learning Features
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
di: Yamashita, Natsuo, et al.
Pubblicazione: (2024)
SELM: Enhancing Speech Emotion Recognition for Out-of-Domain Scenarios
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
di: Bukhari, Hazim, et al.
Pubblicazione: (2024)
Voice-ENHANCE: Speech Restoration using a Diffusion-based Voice Conversion Framework
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
di: Byun, Kyungguen, et al.
Pubblicazione: (2025)
Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
di: Zheng, Xiuwen, et al.
Pubblicazione: (2024)
Quality Assessment of Noisy and Enhanced Speech with Limited Data: UWB-NTIS System for VoiceMOS 2024
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
di: Kunešová, Marie, et al.
Pubblicazione: (2025)
In-Materia Speech Recognition
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
di: Zolfagharinejad, Mohamadreza, et al.
Pubblicazione: (2024)
Amplifying Artifacts with Speech Enhancement in Voice Anti-spoofing
di: Trachu, Thanapat, et al.
Pubblicazione: (2025)
di: Trachu, Thanapat, et al.
Pubblicazione: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
di: Wang, Shiyao, et al.
Pubblicazione: (2025)
SF-Speech: Straightened Flow for Zero-Shot Voice Clone
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
di: Li, Xuyuan, et al.
Pubblicazione: (2024)
The VoiceMOS Challenge 2024: Beyond Speech Quality Prediction
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
di: Huang, Wen-Chin, et al.
Pubblicazione: (2024)
RAVE for Speech: Efficient Voice Conversion at High Sampling Rates
di: Bargum, Anders R., et al.
Pubblicazione: (2024)
di: Bargum, Anders R., et al.
Pubblicazione: (2024)
Enhancing Polyglot Voices by Leveraging Cross-Lingual Fine-Tuning in Any-to-One Voice Conversion
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2024)
di: Ruggiero, Giuseppe, et al.
Pubblicazione: (2024)
EAD-VC: Enhancing Speech Auto-Disentanglement for Voice Conversion with IFUB Estimator and Joint Text-Guided Consistent Learning
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
di: Liang, Ziqi, et al.
Pubblicazione: (2024)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
di: Nasretdinov, Rauf, et al.
Pubblicazione: (2025)
VoiceRestore: Flow-Matching Transformers for Speech Recording Quality Restoration
di: Kirdey, Stanislav
Pubblicazione: (2025)
di: Kirdey, Stanislav
Pubblicazione: (2025)
VC-ENHANCE: Speech Restoration with Integrated Noise Suppression and Voice Conversion
di: Byun, Kyungguen, et al.
Pubblicazione: (2024)
di: Byun, Kyungguen, et al.
Pubblicazione: (2024)
NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple Speakers
di: Park, Nohil, et al.
Pubblicazione: (2024)
di: Park, Nohil, et al.
Pubblicazione: (2024)
Optimal Real-Weighted Beamforming With Application to Linear and Spherical Arrays
di: Tourbabin, V., et al.
Pubblicazione: (2024)
di: Tourbabin, V., et al.
Pubblicazione: (2024)
Speech Emotion Recognition with ASR Integration
di: Li, Yuanchao
Pubblicazione: (2026)
di: Li, Yuanchao
Pubblicazione: (2026)
Zero-Shot Recognition of Dysarthric Speech Using Commercial Automatic Speech Recognition and Multimodal Large Language Models
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
di: Alsayegh, Ali, et al.
Pubblicazione: (2025)
The RoyalFlush Automatic Speech Diarization and Recognition System for In-Car Multi-Channel Automatic Speech Recognition Challenge
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
di: Tian, Jingguang, et al.
Pubblicazione: (2024)
Speech Enhancement with Dual-path Multi-Channel Linear Prediction Filter and Multi-norm Beamforming
di: Qin, Chengyuan, et al.
Pubblicazione: (2025)
di: Qin, Chengyuan, et al.
Pubblicazione: (2025)
Efficient Long-Form Speech Recognition for General Speech In-Context Learning
di: Yen, Hao, et al.
Pubblicazione: (2024)
di: Yen, Hao, et al.
Pubblicazione: (2024)
CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition
di: Wang, He, et al.
Pubblicazione: (2024)
di: Wang, He, et al.
Pubblicazione: (2024)
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
di: Hou, Yixuan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
di: de Groot, Dimme, et al.
Pubblicazione: (2025) -
A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition
di: de Groot, Dimme, et al.
Pubblicazione: (2026) -
Improving Whispered Speech Recognition Performance using Pseudo-whispered based Data Augmentation
di: Lin, Zhaofeng, et al.
Pubblicazione: (2023) -
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
di: Patel, Tanvina, et al.
Pubblicazione: (2025) -
Constant Directivity Loudspeaker Beamforming
di: Luo, Yuancheng
Pubblicazione: (2024)