Saved in:
Bibliographic Details
Main Authors: Li, Zixuan, Zhang, Xueliang, Zhao, Changjiang, Gao, Shuai, Miao, Lei, Yan, Zhipeng, Sun, Ying, Zhu, Chong
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2511.05945
Tags: Add Tag
No Tags, Be the first to tag this record!
Table of Contents:
  • The mean squared error (MSE) is a ubiquitous loss function for speech enhancement, but its problem is that the error cannot reflect the auditory perception quality. This is because MSE causes models to over-emphasize low-frequency components which has high energy, leading to the inadequate modeling of perceptually important high-frequency information. To overcome this limitation, we propose a perceptually-weighted loss function grounded in psychoacoustic principles. Specifically, it leverages equal-loudness contours to assign frequency-dependent weights to the reconstruction error, thereby penalizing deviations in a way aligning with human auditory sensitivity. The proposed loss is model-agnostic and flexible, demonstrating strong generality. Experiments on the VoiceBank+DEMAND dataset show that replacing MSE with our loss in a GTCRN model elevates the WB-PESQ score from 2.17 to 2.93-a significant improvement in perceptual quality.