Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Räsänen, Okko
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912961086357504
author Räsänen, Okko
author_facet Räsänen, Okko
contents Learning to understand speech appears almost effortless for typically developing infants, yet from an information-processing perspective, acquiring a language from acoustic speech is an enormous challenge. This chapter reviews recent developments in using computational models to understand early language acquisition from speech and audiovisual input. The focus is on self-supervised and visually grounded models of perceptual learning. We show how these models are becoming increasingly powerful in learning various aspects of speech without strong linguistic priors, and how many features of early language development can be explained through a shared set of learning principles-principles broadly compatible with multiple theories of language acquisition and human cognition. We also discuss how modern learning simulations are gradually becoming more realistic, both in terms of input data and in linking model behavior to empirical findings on infant language development.
format Preprint
id arxiv_https___arxiv_org_abs_2603_08359
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
Räsänen, Okko
Computation and Language
Artificial Intelligence
Audio and Speech Processing
I.2.6; I.2.7; J.4
Learning to understand speech appears almost effortless for typically developing infants, yet from an information-processing perspective, acquiring a language from acoustic speech is an enormous challenge. This chapter reviews recent developments in using computational models to understand early language acquisition from speech and audiovisual input. The focus is on self-supervised and visually grounded models of perceptual learning. We show how these models are becoming increasingly powerful in learning various aspects of speech without strong linguistic priors, and how many features of early language development can be explained through a shared set of learning principles-principles broadly compatible with multiple theories of language acquisition and human cognition. We also discuss how modern learning simulations are gradually becoming more realistic, both in terms of input data and in linking model behavior to empirical findings on infant language development.
title Computational modeling of early language learning from acoustic speech and audiovisual input without linguistic priors
topic Computation and Language
Artificial Intelligence
Audio and Speech Processing
I.2.6; I.2.7; J.4
url https://arxiv.org/abs/2603.08359