Unsupervised Variational Acoustic Clustering
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866912836068835328 |
|---|---|
| author | Fiorio, Luan Vinícius Defraene, Bruno David, Johan Widdershoven, Frans van Houtum, Wim Aarts, Ronald M. |
| author_facet | Fiorio, Luan Vinícius Defraene, Bruno David, Johan Widdershoven, Frans van Houtum, Wim Aarts, Ronald M. |
| contents | We propose an unsupervised variational acoustic clustering model for clustering audio data in the time-frequency domain. The model leverages variational inference, extended to an autoencoder framework, with a Gaussian mixture model as a prior for the latent space. Specifically designed for audio applications, we introduce a convolutional-recurrent variational autoencoder optimized for efficient time-frequency processing. Our experimental results considering a spoken digits dataset demonstrate a significant improvement in accuracy and clustering performance compared to traditional methods, showcasing the model's enhanced ability to capture complex audio patterns. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2503_18579 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Unsupervised Variational Acoustic Clustering Fiorio, Luan Vinícius Defraene, Bruno David, Johan Widdershoven, Frans van Houtum, Wim Aarts, Ronald M. Audio and Speech Processing Signal Processing We propose an unsupervised variational acoustic clustering model for clustering audio data in the time-frequency domain. The model leverages variational inference, extended to an autoencoder framework, with a Gaussian mixture model as a prior for the latent space. Specifically designed for audio applications, we introduce a convolutional-recurrent variational autoencoder optimized for efficient time-frequency processing. Our experimental results considering a spoken digits dataset demonstrate a significant improvement in accuracy and clustering performance compared to traditional methods, showcasing the model's enhanced ability to capture complex audio patterns. |
| title | Unsupervised Variational Acoustic Clustering |
| topic | Audio and Speech Processing Signal Processing |
| url | https://arxiv.org/abs/2503.18579 |