Morpheme Induction for Emergent Language

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Boldt, Brendon, Mortensen, David
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915532658180096
author Boldt, Brendon
Mortensen, David
author_facet Boldt, Brendon
Mortensen, David
contents We introduce CSAR, an algorithm for inducing morphemes from emergent language corpora of parallel utterances and meanings. It is a greedy algorithm that (1) weights morphemes based on mutual information between forms and meanings, (2) selects the highest-weighted pair, (3) removes it from the corpus, and (4) repeats the process to induce further morphemes (i.e., Count, Select, Ablate, Repeat). The effectiveness of CSAR is first validated on procedurally generated datasets and compared against baselines for related tasks. Second, we validate CSAR's performance on human language data to show that the algorithm makes reasonable predictions in adjacent domains. Finally, we analyze a handful of emergent languages, quantifying linguistic characteristics like degree of synonymy and polysemy.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Morpheme Induction for Emergent Language
Boldt, Brendon
Mortensen, David
Computation and Language
I.2.7; I.6.m
We introduce CSAR, an algorithm for inducing morphemes from emergent language corpora of parallel utterances and meanings. It is a greedy algorithm that (1) weights morphemes based on mutual information between forms and meanings, (2) selects the highest-weighted pair, (3) removes it from the corpus, and (4) repeats the process to induce further morphemes (i.e., Count, Select, Ablate, Repeat). The effectiveness of CSAR is first validated on procedurally generated datasets and compared against baselines for related tasks. Second, we validate CSAR's performance on human language data to show that the algorithm makes reasonable predictions in adjacent domains. Finally, we analyze a handful of emergent languages, quantifying linguistic characteristics like degree of synonymy and polysemy.
title Morpheme Induction for Emergent Language
topic Computation and Language
I.2.7; I.6.m
url https://arxiv.org/abs/2510.03439