Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Berger, Clémentine, Badeau, Roland, Essid, Slim
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915171240247296
author Berger, Clémentine
Badeau, Roland
Essid, Slim
author_facet Berger, Clémentine
Badeau, Roland
Essid, Slim
contents People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise's frequency components due to the effect of simultaneous masking. In this article, we propose a neural network based on a psychoacoustic masking model, designed to enhance the music's ability to mask ambient noise by reshaping its spectral envelope with predicted filter frequency responses. The model is trained with a perceptual loss function that balances two constraints: effectively masking the noise while preserving the original music mix and the user's chosen listening level. We evaluate our approach on simulated data replicating a user's experience of listening to music with headphones in a noisy environment. The results, based on defined objective metrics, demonstrate that our system improves the state of the art.
format Preprint
id arxiv_https___arxiv_org_abs_2502_17527
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
Berger, Clémentine
Badeau, Roland
Essid, Slim
Sound
Artificial Intelligence
Audio and Speech Processing
Signal Processing
People often listen to music in noisy environments, seeking to isolate themselves from ambient sounds. Indeed, a music signal can mask some of the noise's frequency components due to the effect of simultaneous masking. In this article, we propose a neural network based on a psychoacoustic masking model, designed to enhance the music's ability to mask ambient noise by reshaping its spectral envelope with predicted filter frequency responses. The model is trained with a perceptual loss function that balances two constraints: effectively masking the noise while preserving the original music mix and the user's chosen listening level. We evaluate our approach on simulated data replicating a user's experience of listening to music with headphones in a noisy environment. The results, based on defined objective metrics, demonstrate that our system improves the state of the art.
title Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
topic Sound
Artificial Intelligence
Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2502.17527