Revealing the Role of Audio Channels in ASR Performance Degradation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Kuan-Tang, Chen, Li-Wei, Lee, Hung-Shin, Chen, Berlin, Wang, Hsin-Min
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915456467599360
author Huang, Kuan-Tang
Chen, Li-Wei
Lee, Hung-Shin
Chen, Berlin
Wang, Hsin-Min
author_facet Huang, Kuan-Tang
Chen, Li-Wei
Lee, Hung-Shin
Chen, Berlin
Wang, Hsin-Min
contents Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.
format Preprint
id arxiv_https___arxiv_org_abs_2508_08967
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Revealing the Role of Audio Channels in ASR Performance Degradation
Huang, Kuan-Tang
Chen, Li-Wei
Lee, Hung-Shin
Chen, Berlin
Wang, Hsin-Min
Sound
Artificial Intelligence
Computation and Language
Pre-trained automatic speech recognition (ASR) models have demonstrated strong performance on a variety of tasks. However, their performance can degrade substantially when the input audio comes from different recording channels. While previous studies have demonstrated this phenomenon, it is often attributed to the mismatch between training and testing corpora. This study argues that variations in speech characteristics caused by different recording channels can fundamentally harm ASR performance. To address this limitation, we propose a normalization technique designed to mitigate the impact of channel variation by aligning internal feature representations in the ASR model with those derived from a clean reference channel. This approach significantly improves ASR performance on previously unseen channels and languages, highlighting its ability to generalize across channel and language differences.
title Revealing the Role of Audio Channels in ASR Performance Degradation
topic Sound
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2508.08967