Data Efficient Child-Adult Speaker Diarization with Simulated Conversations

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Anfeng, Feng, Tiantian, Tager-Flusberg, Helen, Lord, Catherine, Narayanan, Shrikanth
Formato: Preprint
Publicado: 2024
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913889739866112
author Xu, Anfeng
Feng, Tiantian
Tager-Flusberg, Helen
Lord, Catherine
Narayanan, Shrikanth
author_facet Xu, Anfeng
Feng, Tiantian
Tager-Flusberg, Helen
Lord, Catherine
Narayanan, Shrikanth
contents Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2409_08881
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
Xu, Anfeng
Feng, Tiantian
Tager-Flusberg, Helen
Lord, Catherine
Narayanan, Shrikanth
Audio and Speech Processing
Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available.
title Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
topic Audio and Speech Processing
url https://arxiv.org/abs/2409.08881