Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
Fuente:
arXiv
Guardado en:
| Autores principales: | Ochiai, Tsubasa, Iwamoto, Kazuma, Delcroix, Marc, Ikeshita, Rintaro, Sato, Hiroshi, Araki, Shoko, Katagiri, Shigeru |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Target Speech Extraction with Pre-trained Self-supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2024)
por: Peng, Junyi, et al.
Publicado: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2025)
por: Peng, Junyi, et al.
Publicado: (2025)
Probing Self-supervised Learning Models with Target Speech Extraction
por: Peng, Junyi, et al.
Publicado: (2024)
por: Peng, Junyi, et al.
Publicado: (2024)
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
por: Lohmann, Anselm, et al.
Publicado: (2025)
por: Lohmann, Anselm, et al.
Publicado: (2025)
Frontend Token Enhancement for Token-Based Speech Recognition
por: Ashihara, Takanori, et al.
Publicado: (2026)
por: Ashihara, Takanori, et al.
Publicado: (2026)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
por: Sato, Hiroshi, et al.
Publicado: (2025)
por: Sato, Hiroshi, et al.
Publicado: (2025)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
por: Tammen, Marvin, et al.
Publicado: (2024)
por: Tammen, Marvin, et al.
Publicado: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
por: Ashihara, Takanori, et al.
Publicado: (2024)
por: Ashihara, Takanori, et al.
Publicado: (2024)
Interaural time difference loss for binaural target sound extraction
por: Hernandez-Olivan, Carlos, et al.
Publicado: (2024)
por: Hernandez-Olivan, Carlos, et al.
Publicado: (2024)
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
por: Hernandez-Olivan, Carlos, et al.
Publicado: (2024)
por: Hernandez-Olivan, Carlos, et al.
Publicado: (2024)
Mamba-based Segmentation Model for Speaker Diarization
por: Plaquet, Alexis, et al.
Publicado: (2024)
por: Plaquet, Alexis, et al.
Publicado: (2024)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
por: von Neumann, Thilo, et al.
Publicado: (2023)
por: von Neumann, Thilo, et al.
Publicado: (2023)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
por: Plaquet, Alexis, et al.
Publicado: (2025)
por: Plaquet, Alexis, et al.
Publicado: (2025)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
por: Sato, Hiroshi, et al.
Publicado: (2024)
por: Sato, Hiroshi, et al.
Publicado: (2024)
Robust Speech Recognition with Schrödinger Bridge-Based Speech Enhancement
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
por: Nasretdinov, Rauf, et al.
Publicado: (2025)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
por: Han, Seungu, et al.
Publicado: (2026)
por: Han, Seungu, et al.
Publicado: (2026)
Unified Architecture and Unsupervised Speech Disentanglement for Speaker Embedding-Free Enrollment in Personalized Speech Enhancement
por: Huang, Ziling, et al.
Publicado: (2025)
por: Huang, Ziling, et al.
Publicado: (2025)
Geneses: Unified Generative Speech Enhancement and Separation
por: Asai, Kohei, et al.
Publicado: (2026)
por: Asai, Kohei, et al.
Publicado: (2026)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
por: Chen, Shuangyuan, et al.
Publicado: (2025)
por: Chen, Shuangyuan, et al.
Publicado: (2025)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
por: Pálka, Petr, et al.
Publicado: (2024)
por: Pálka, Petr, et al.
Publicado: (2024)
Retrieval Augmented Correction of Named Entity Speech Recognition Errors
por: Pusateri, Ernest, et al.
Publicado: (2024)
por: Pusateri, Ernest, et al.
Publicado: (2024)
SuPseudo: A Pseudo-supervised Learning Method for Neural Speech Enhancement in Far-field Speech Recognition
por: Luo, Longjie, et al.
Publicado: (2025)
por: Luo, Longjie, et al.
Publicado: (2025)
FlexIO: Flexible Single- and Multi-Channel Speech Separation and Enhancement
por: Masuyama, Yoshiki, et al.
Publicado: (2025)
por: Masuyama, Yoshiki, et al.
Publicado: (2025)
Latent-Level Enhancement with Flow Matching for Robust Automatic Speech Recognition
por: Yang, Da-Hee, et al.
Publicado: (2026)
por: Yang, Da-Hee, et al.
Publicado: (2026)
MOVER: Combining Multiple Meeting Recognition Systems
por: Kamo, Naoyuki, et al.
Publicado: (2025)
por: Kamo, Naoyuki, et al.
Publicado: (2025)
Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module
por: Cui, Zhongjian, et al.
Publicado: (2025)
por: Cui, Zhongjian, et al.
Publicado: (2025)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
por: Pandey, Ashutosh, et al.
Publicado: (2024)
por: Pandey, Ashutosh, et al.
Publicado: (2024)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
por: Tsunoo, Emiru, et al.
Publicado: (2024)
por: Tsunoo, Emiru, et al.
Publicado: (2024)
In-Materia Speech Recognition
por: Zolfagharinejad, Mohamadreza, et al.
Publicado: (2024)
por: Zolfagharinejad, Mohamadreza, et al.
Publicado: (2024)
Plugin Speech Enhancement: A Universal Speech Enhancement Framework Inspired by Dynamic Neural Network
por: Chen, Yanan, et al.
Publicado: (2024)
por: Chen, Yanan, et al.
Publicado: (2024)
Learning Expressive Disentangled Speech Representations with Soft Speech Units and Adversarial Style Augmentation
por: Deng, Yimin, et al.
Publicado: (2024)
por: Deng, Yimin, et al.
Publicado: (2024)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
por: Horiguchi, Shota, et al.
Publicado: (2024)
por: Horiguchi, Shota, et al.
Publicado: (2024)
Guided Speaker Embedding
por: Horiguchi, Shota, et al.
Publicado: (2024)
por: Horiguchi, Shota, et al.
Publicado: (2024)
Advancing Electrolaryngeal Speech Enhancement Through Speech-Text Representation Learning
por: Ma, Ding, et al.
Publicado: (2026)
por: Ma, Ding, et al.
Publicado: (2026)
Objective and Subjective Evaluation of Diffusion-Based Speech Enhancement for Dysarthric Speech
por: de Groot, Dimme, et al.
Publicado: (2025)
por: de Groot, Dimme, et al.
Publicado: (2025)
Bridging the Gap: Integrating Pre-trained Speech Enhancement and Recognition Models for Robust Speech Recognition
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
por: Wang, Kuan-Chen, et al.
Publicado: (2024)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
por: Ashihara, Takanori, et al.
Publicado: (2023)
por: Ashihara, Takanori, et al.
Publicado: (2023)
Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model
por: Li, Jizhen, et al.
Publicado: (2025)
por: Li, Jizhen, et al.
Publicado: (2025)
Absorbing Discrete Diffusion for Speech Enhancement
por: Gonzalez, Philippe
Publicado: (2026)
por: Gonzalez, Philippe
Publicado: (2026)
Ejemplares similares
-
Target Speech Extraction with Pre-trained Self-supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2024) -
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2025) -
Probing Self-supervised Learning Models with Target Speech Extraction
por: Peng, Junyi, et al.
Publicado: (2024) -
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
por: Lohmann, Anselm, et al.
Publicado: (2025) -
Frontend Token Enhancement for Token-Based Speech Recognition
por: Ashihara, Takanori, et al.
Publicado: (2026)