Unsupervised Variational Acoustic Clustering

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Fiorio, Luan Vinícius, Defraene, Bruno, David, Johan, Widdershoven, Frans, van Houtum, Wim, Aarts, Ronald M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912836068835328
author Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
author_facet Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
contents We propose an unsupervised variational acoustic clustering model for clustering audio data in the time-frequency domain. The model leverages variational inference, extended to an autoencoder framework, with a Gaussian mixture model as a prior for the latent space. Specifically designed for audio applications, we introduce a convolutional-recurrent variational autoencoder optimized for efficient time-frequency processing. Our experimental results considering a spoken digits dataset demonstrate a significant improvement in accuracy and clustering performance compared to traditional methods, showcasing the model's enhanced ability to capture complex audio patterns.
format Preprint
id arxiv_https___arxiv_org_abs_2503_18579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unsupervised Variational Acoustic Clustering
Fiorio, Luan Vinícius
Defraene, Bruno
David, Johan
Widdershoven, Frans
van Houtum, Wim
Aarts, Ronald M.
Audio and Speech Processing
Signal Processing
We propose an unsupervised variational acoustic clustering model for clustering audio data in the time-frequency domain. The model leverages variational inference, extended to an autoencoder framework, with a Gaussian mixture model as a prior for the latent space. Specifically designed for audio applications, we introduce a convolutional-recurrent variational autoencoder optimized for efficient time-frequency processing. Our experimental results considering a spoken digits dataset demonstrate a significant improvement in accuracy and clustering performance compared to traditional methods, showcasing the model's enhanced ability to capture complex audio patterns.
title Unsupervised Variational Acoustic Clustering
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2503.18579