voc2vec: A Foundation Model for Non-Verbal Vocalization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Koudounas, Alkis, La Quatra, Moreno, Siniscalchi, Marco Sabato, Baralis, Elena
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916626053464064
author Koudounas, Alkis
La Quatra, Moreno
Siniscalchi, Marco Sabato
Baralis, Elena
author_facet Koudounas, Alkis
La Quatra, Moreno
Siniscalchi, Marco Sabato
Baralis, Elena
contents Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various real-world applications. Audio foundation models well handle non-speech data but also fail to capture the nuanced features of non-verbal human sounds. In this work, we aim to overcome the above shortcoming and propose a novel foundation model, termed voc2vec, specifically designed for non-verbal human data leveraging exclusively open-source non-verbal audio datasets. We employ a collection of 10 datasets covering around 125 hours of non-verbal audio. Experimental results prove that voc2vec is effective in non-verbal vocalization classification, and it outperforms conventional speech and audio foundation models. Moreover, voc2vec consistently outperforms strong baselines, namely OpenSmile and emotion2vec, on six different benchmark datasets. To the best of the authors' knowledge, voc2vec is the first universal representation model for vocalization tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_16298
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle voc2vec: A Foundation Model for Non-Verbal Vocalization
Koudounas, Alkis
La Quatra, Moreno
Siniscalchi, Marco Sabato
Baralis, Elena
Audio and Speech Processing
Sound
Speech foundation models have demonstrated exceptional capabilities in speech-related tasks. Nevertheless, these models often struggle with non-verbal audio data, such as vocalizations, baby crying, etc., which are critical for various real-world applications. Audio foundation models well handle non-speech data but also fail to capture the nuanced features of non-verbal human sounds. In this work, we aim to overcome the above shortcoming and propose a novel foundation model, termed voc2vec, specifically designed for non-verbal human data leveraging exclusively open-source non-verbal audio datasets. We employ a collection of 10 datasets covering around 125 hours of non-verbal audio. Experimental results prove that voc2vec is effective in non-verbal vocalization classification, and it outperforms conventional speech and audio foundation models. Moreover, voc2vec consistently outperforms strong baselines, namely OpenSmile and emotion2vec, on six different benchmark datasets. To the best of the authors' knowledge, voc2vec is the first universal representation model for vocalization tasks.
title voc2vec: A Foundation Model for Non-Verbal Vocalization
topic Audio and Speech Processing
Sound
url https://arxiv.org/abs/2502.16298