An open-source voice type classifier for child-centered daylong recordings

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Lavechin, Marvin, Bousbib, Ruben, Bredin, Hervé, Dupoux, Emmanuel, Cristia, Alejandrina
Format: Preprint
Publié: 2020
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917950099816448
author Lavechin, Marvin
Bousbib, Ruben
Bredin, Hervé
Dupoux, Emmanuel
Cristia, Alejandrina
author_facet Lavechin, Marvin
Bousbib, Ruben
Bredin, Hervé
Dupoux, Emmanuel
Cristia, Alejandrina
contents Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a wide variety of conditions would be particularly useful for language acquisition studies in which researchers are interested in the quantity and quality of the speech that children hear and produce, as well as for early diagnosis and measuring effects of remediation. In this paper, we present our approach to designing an open-source neural network to classify audio segments into vocalizations produced by the child wearing the recording device, vocalizations produced by other children, adult male speech, and adult female speech. To this end, we gathered diverse child-centered corpora which sums up to a total of 260 hours of recordings and covers 10 languages. Our model can be used as input for downstream tasks such as estimating the number of words produced by adult speakers, or the number of linguistic units produced by children. Our architecture combines SincNet filters with a stack of recurrent layers and outperforms by a large margin the state-of-the-art system, the Language ENvironment Analysis (LENA) that has been used in numerous child language studies.
format Preprint
id arxiv_https___arxiv_org_abs_2005_12656
institution arXiv
publishDate 2020
record_format arxiv
spellingShingle An open-source voice type classifier for child-centered daylong recordings
Lavechin, Marvin
Bousbib, Ruben
Bredin, Hervé
Dupoux, Emmanuel
Cristia, Alejandrina
Audio and Speech Processing
I.2.7
Spontaneous conversations in real-world settings such as those found in child-centered recordings have been shown to be amongst the most challenging audio files to process. Nevertheless, building speech processing models handling such a wide variety of conditions would be particularly useful for language acquisition studies in which researchers are interested in the quantity and quality of the speech that children hear and produce, as well as for early diagnosis and measuring effects of remediation. In this paper, we present our approach to designing an open-source neural network to classify audio segments into vocalizations produced by the child wearing the recording device, vocalizations produced by other children, adult male speech, and adult female speech. To this end, we gathered diverse child-centered corpora which sums up to a total of 260 hours of recordings and covers 10 languages. Our model can be used as input for downstream tasks such as estimating the number of words produced by adult speakers, or the number of linguistic units produced by children. Our architecture combines SincNet filters with a stack of recurrent layers and outperforms by a large margin the state-of-the-art system, the Language ENvironment Analysis (LENA) that has been used in numerous child language studies.
title An open-source voice type classifier for child-centered daylong recordings
topic Audio and Speech Processing
I.2.7
url https://arxiv.org/abs/2005.12656