Context-aware child-directed speech detection from long-form recordings

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Charlot, Théo, Kunze, Tarek, Sheth, Kaveri K., Cristia, Alejandrina, Lavechin, Marvin
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910277272862720
author Charlot, Théo
Kunze, Tarek
Sheth, Kaveri K.
Cristia, Alejandrina
Lavechin, Marvin
author_facet Charlot, Théo
Kunze, Tarek
Sheth, Kaveri K.
Cristia, Alejandrina
Lavechin, Marvin
contents Automatically distinguishing child-directed speech from adult-directed speech in long-form recordings is key to scalable analyses of children's language environments. Existing approaches process utterances in isolation and have been evaluated primarily on English. We address these gaps along three dimensions. First, we fine-tune and evaluate six-self supervised models on a multilingual dataset of 182 children, showing that in-domain pre-training on child-centered recordings substantially outperforms models trained on adult speech. Second, we demonstrate that incorporating surrounding context substantially improves classification, with an absolute gain of 13.8% in average F1-score. Third, we evaluate our model in a realistic end-to-end pipeline, from adult speech detection to addressee classification, showing that performance drops under automatic segmentation but still consistently outperforms a rule-based baseline.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01134
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Context-aware child-directed speech detection from long-form recordings
Charlot, Théo
Kunze, Tarek
Sheth, Kaveri K.
Cristia, Alejandrina
Lavechin, Marvin
Audio and Speech Processing
Machine Learning
Sound
Automatically distinguishing child-directed speech from adult-directed speech in long-form recordings is key to scalable analyses of children's language environments. Existing approaches process utterances in isolation and have been evaluated primarily on English. We address these gaps along three dimensions. First, we fine-tune and evaluate six-self supervised models on a multilingual dataset of 182 children, showing that in-domain pre-training on child-centered recordings substantially outperforms models trained on adult speech. Second, we demonstrate that incorporating surrounding context substantially improves classification, with an absolute gain of 13.8% in average F1-score. Third, we evaluate our model in a realistic end-to-end pipeline, from adult speech detection to addressee classification, showing that performance drops under automatic segmentation but still consistently outperforms a rule-based baseline.
title Context-aware child-directed speech detection from long-form recordings
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2606.01134