Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Moritz, Martin, Olán, Toni, Virtanen, Tuomas
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913570138095616
author Moritz, Martin
Olán, Toni
Virtanen, Tuomas
author_facet Moritz, Martin
Olán, Toni
Virtanen, Tuomas
contents Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to be used in deep neural network based audio watermarking systems. The loss is based on the noise-to-mask ratio (NMR), which is a model of the psychoacoustic masking effect characteristic of the human ear. We use the NMR loss between marked and host signals to train the deep neural models and we evaluate the objective quality with PEAQ and the subjective quality with a MUSHRA test. Both objective and subjective tests show that models trained with NMR loss generate more transparent watermarks than models trained with the conventionally used MSE loss
format Preprint
id arxiv_https___arxiv_org_abs_2408_15553
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
Moritz, Martin
Olán, Toni
Virtanen, Tuomas
Audio and Speech Processing
Sound
Digital audio watermarking consists in inserting a message into audio signals in a transparent way and can be used to allow automatic recognition of audio material and management of the copyrights. We propose a perceptual loss function to be used in deep neural network based audio watermarking systems. The loss is based on the noise-to-mask ratio (NMR), which is a model of the psychoacoustic masking effect characteristic of the human ear. We use the NMR loss between marked and host signals to train the deep neural models and we evaluate the objective quality with PEAQ and the subjective quality with a MUSHRA test. Both objective and subjective tests show that models trained with NMR loss generate more transparent watermarks than models trained with the conventionally used MSE loss
title Noise-to-mask Ratio Loss for Deep Neural Network based Audio Watermarking
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2408.15553