Linguistic Structure from a Bottleneck on Sequential Information Processing

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Futrell, Richard, Hahn, Michael
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912715314823168
author Futrell, Richard
Hahn, Michael
author_facet Futrell, Richard
Hahn, Michael
contents Human language has a distinct systematic structure, where utterances break into individually meaningful words which are combined to form phrases. We show that natural-language-like systematicity arises in codes that are constrained by a statistical measure of complexity called predictive information, also known as excess entropy. Predictive information is the mutual information between the past and future of a stochastic process. In simulations, we find that such codes break messages into groups of approximately independent features which are expressed systematically and locally, corresponding to words and phrases. Next, drawing on crosslinguistic text corpora, we find that actual human languages are structured in a way that reduces predictive information compared to baselines at the levels of phonology, morphology, syntax, and lexical semantics. Our results establish a link between the statistical and algebraic structure of language and reinforce the idea that these structures are shaped by communication under general cognitive constraints.
format Preprint
id arxiv_https___arxiv_org_abs_2405_12109
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Linguistic Structure from a Bottleneck on Sequential Information Processing
Futrell, Richard
Hahn, Michael
Computation and Language
Information Theory
Human language has a distinct systematic structure, where utterances break into individually meaningful words which are combined to form phrases. We show that natural-language-like systematicity arises in codes that are constrained by a statistical measure of complexity called predictive information, also known as excess entropy. Predictive information is the mutual information between the past and future of a stochastic process. In simulations, we find that such codes break messages into groups of approximately independent features which are expressed systematically and locally, corresponding to words and phrases. Next, drawing on crosslinguistic text corpora, we find that actual human languages are structured in a way that reduces predictive information compared to baselines at the levels of phonology, morphology, syntax, and lexical semantics. Our results establish a link between the statistical and algebraic structure of language and reinforce the idea that these structures are shaped by communication under general cognitive constraints.
title Linguistic Structure from a Bottleneck on Sequential Information Processing
topic Computation and Language
Information Theory
url https://arxiv.org/abs/2405.12109