Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bjare, Mathias Rose, Cantisani, Giorgia, Pasini, Marco, Lattner, Stefan, Widmer, Gerhard
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915608492244992
author Bjare, Mathias Rose
Cantisani, Giorgia
Pasini, Marco
Lattner, Stefan
Widmer, Gerhard
author_facet Bjare, Mathias Rose
Cantisani, Giorgia
Pasini, Marco
Lattner, Stefan
Widmer, Gerhard
contents We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptual losses, yields encodings that are structured according to a perceptual hierarchy. We demonstrate the emergence of this hierarchical structure by showing that, after training an audio autoencoder in this manner, perceptually salient information is captured in coarser representation structures than with conventional training. Furthermore, we show that such perceptual hierarchies improve latent diffusion decoding in the context of estimating surprisal in music pitches and predicting EEG-brain responses to music listening. Pretrained weights are available on github.com/CPJKU/pa-audioic.
format Preprint
id arxiv_https___arxiv_org_abs_2511_05350
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
Bjare, Mathias Rose
Cantisani, Giorgia
Pasini, Marco
Lattner, Stefan
Widmer, Gerhard
Sound
Artificial Intelligence
We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptual losses, yields encodings that are structured according to a perceptual hierarchy. We demonstrate the emergence of this hierarchical structure by showing that, after training an audio autoencoder in this manner, perceptually salient information is captured in coarser representation structures than with conventional training. Furthermore, we show that such perceptual hierarchies improve latent diffusion decoding in the context of estimating surprisal in music pitches and predicting EEG-brain responses to music listening. Pretrained weights are available on github.com/CPJKU/pa-audioic.
title Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2511.05350