Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Räsänen, Okko, Kocharov, Daniil
Format:	Preprint
Published:	2024
Subjects:	Computation and Language Sound Audio and Speech Processing
Online Access:	https://arxiv.org/abs/2405.07700
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866911875489333248
author	Räsänen, Okko Kocharov, Daniil
author_facet	Räsänen, Okko Kocharov, Daniil
contents	Child-directed speech (CDS) is a particular type of speech that adults use when addressing young children. Its properties also change as a function of extralinguistic factors, such as age of the child being addressed. Access to large amounts of representative and varied CDS would be useful for child language research, as this would enable controlled computational modeling experiments of infant language acquisition with realistic input in terms of quality and quantity. In this study, we describe an approach to model age-dependent linguistic properties of CDS using a language model (LM) trained on CDS transcripts and ages of the recipient children, as obtained from North American English corpora of the CHILDES database. The created LM can then be used to stochastically generate synthetic CDS transcripts in an age-appropriate manner, thereby scaling beyond the original datasets in size. We compare characteristics of the generated CDS against the real speech addressed at children of different ages, showing that the LM manages to capture age-dependent changes in CDS, except for a slight difference in the effective vocabulary size. As a side product, we also provide a systematic characterization of age-dependent linguistic properties of CDS in CHILDES, illustrating how all measured aspects of the CDS change with children's age.
format	Preprint
id	arxiv_https___arxiv_org_abs_2405_07700
institution	arXiv
publishDate	2024
record_format	arxiv
spellingShingle	Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech Räsänen, Okko Kocharov, Daniil Computation and Language Sound Audio and Speech Processing Child-directed speech (CDS) is a particular type of speech that adults use when addressing young children. Its properties also change as a function of extralinguistic factors, such as age of the child being addressed. Access to large amounts of representative and varied CDS would be useful for child language research, as this would enable controlled computational modeling experiments of infant language acquisition with realistic input in terms of quality and quantity. In this study, we describe an approach to model age-dependent linguistic properties of CDS using a language model (LM) trained on CDS transcripts and ages of the recipient children, as obtained from North American English corpora of the CHILDES database. The created LM can then be used to stochastically generate synthetic CDS transcripts in an age-appropriate manner, thereby scaling beyond the original datasets in size. We compare characteristics of the generated CDS against the real speech addressed at children of different ages, showing that the LM manages to capture age-dependent changes in CDS, except for a slight difference in the effective vocabulary size. As a side product, we also provide a systematic characterization of age-dependent linguistic properties of CDS in CHILDES, illustrating how all measured aspects of the CDS change with children's age.
title	Age-Dependent Analysis and Stochastic Generation of Child-Directed Speech
topic	Computation and Language Sound Audio and Speech Processing
url	https://arxiv.org/abs/2405.07700

Similar Items