On feature representations for marmoset vocal communication analysis

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sarkar, Eklavya, Wierucka, Kaja, Bosshard, Alexandra B., Burkart, Judith, -Doss, Mathew Magimai.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917992760082432
author Sarkar, Eklavya
Wierucka, Kaja
Bosshard, Alexandra B.
Burkart, Judith
-Doss, Mathew Magimai.
author_facet Sarkar, Eklavya
Wierucka, Kaja
Bosshard, Alexandra B.
Burkart, Judith
-Doss, Mathew Magimai.
contents The acoustic analysis of marmoset (Callithrix jacchus) vocalizations is often used to understand the evolutionary origins of human language. Currently, the analysis is largely carried out in a manual or semi-manual manner. Thus, there is a need to develop automatic call analysis methods. In that direction, research has been limited to the development of analysis methods with small amounts of data or for specific scenarios. Furthermore, there is lack of prior knowledge about what type of information is relevant for different call analysis tasks. To address these issues, as a first step, this paper explores different feature representation methods, namely, HCTSA-based hand-crafted features Catch22, pre-trained self supervised learning (SSL) based features extracted from neural networks trained on human speech and end-to-end acoustic modeling for call-type classification, caller identification and caller sex identification. Through an investigation on three different marmoset call datasets, we demonstrate that SSL-based feature representations and end-to-end acoustic modeling tend to lead to better systems than Catch22 features for call-type and caller classification. Furthermore, we also highlight the impact of signal bandwidth on the obtained task performances.
format Preprint
id arxiv_https___arxiv_org_abs_2504_14981
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle On feature representations for marmoset vocal communication analysis
Sarkar, Eklavya
Wierucka, Kaja
Bosshard, Alexandra B.
Burkart, Judith
-Doss, Mathew Magimai.
Audio and Speech Processing
The acoustic analysis of marmoset (Callithrix jacchus) vocalizations is often used to understand the evolutionary origins of human language. Currently, the analysis is largely carried out in a manual or semi-manual manner. Thus, there is a need to develop automatic call analysis methods. In that direction, research has been limited to the development of analysis methods with small amounts of data or for specific scenarios. Furthermore, there is lack of prior knowledge about what type of information is relevant for different call analysis tasks. To address these issues, as a first step, this paper explores different feature representation methods, namely, HCTSA-based hand-crafted features Catch22, pre-trained self supervised learning (SSL) based features extracted from neural networks trained on human speech and end-to-end acoustic modeling for call-type classification, caller identification and caller sex identification. Through an investigation on three different marmoset call datasets, we demonstrate that SSL-based feature representations and end-to-end acoustic modeling tend to lead to better systems than Catch22 features for call-type and caller classification. Furthermore, we also highlight the impact of signal bandwidth on the obtained task performances.
title On feature representations for marmoset vocal communication analysis
topic Audio and Speech Processing
url https://arxiv.org/abs/2504.14981