Data Efficient Child-Adult Speaker Diarization with Simulated Conversations
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866913889739866112 |
|---|---|
| author | Xu, Anfeng Feng, Tiantian Tager-Flusberg, Helen Lord, Catherine Narayanan, Shrikanth |
| author_facet | Xu, Anfeng Feng, Tiantian Tager-Flusberg, Helen Lord, Catherine Narayanan, Shrikanth |
| contents | Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2409_08881 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Data Efficient Child-Adult Speaker Diarization with Simulated Conversations Xu, Anfeng Feng, Tiantian Tager-Flusberg, Helen Lord, Catherine Narayanan, Shrikanth Audio and Speech Processing Automating child speech analysis is crucial for applications such as neurocognitive assessments. Speaker diarization, which identifies ``who spoke when'', is an essential component of the automated analysis. However, publicly available child-adult speaker diarization solutions are scarce due to privacy concerns and a lack of annotated datasets, while manually annotating data for each scenario is both time-consuming and costly. To overcome these challenges, we propose a data-efficient solution by creating simulated child-adult conversations using AudioSet. We then train a Whisper Encoder-based model, achieving strong zero-shot performance on child-adult speaker diarization using real datasets. The model performance improves substantially when fine-tuned with only 30 minutes of real train data, with LoRA further improving the transfer learning performance. The source code and the child-adult speaker diarization model trained on simulated conversations are publicly available. |
| title | Data Efficient Child-Adult Speaker Diarization with Simulated Conversations |
| topic | Audio and Speech Processing |
| url | https://arxiv.org/abs/2409.08881 |