Internformat: :: Library Catalog

Gespeichert in:

Bibliographische Detailangaben
Hauptverfasser:	Szubert, Ida, Abend, Omri, Schneider, Nathan, Gibbon, Samuel, Mahon, Louis, Goldwater, Sharon, Steedman, Mark
Format:	Preprint
Veröffentlicht:	2021
Schlagworte:	Computation and Language
Online-Zugang:	https://arxiv.org/abs/2109.10952
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

_version_	1866914714910457856
author	Szubert, Ida Abend, Omri Schneider, Nathan Gibbon, Samuel Mahon, Louis Goldwater, Sharon Steedman, Mark
author_facet	Szubert, Ida Abend, Omri Schneider, Nathan Gibbon, Samuel Mahon, Louis Goldwater, Sharon Steedman, Mark
contents	This paper proposes a methodology for constructing such corpora of child directed speech (CDS) paired with sentential logical forms, and uses this method to create two such corpora, in English and Hebrew. The approach enforces a cross-linguistically consistent representation, building on recent advances in dependency representation and semantic parsing. Specifically, the approach involves two steps. First, we annotate the corpora using the Universal Dependencies (UD) scheme for syntactic annotation, which has been developed to apply consistently to a wide variety of domains and typologically diverse languages. Next, we further annotate these data by applying an automatic method for transducing sentential logical forms (LFs) from UD structures. The UD and LF representations have complementary strengths: UD structures are language-neutral and support consistent and reliable annotation by multiple annotators, whereas LFs are neutral as to their syntactic derivation and transparently encode semantic relations. Using this approach, we provide syntactic and semantic annotation for two corpora from CHILDES: Brown's Adam corpus (English; we annotate ~80% of its child-directed utterances), all child-directed utterances from Berman's Hagar corpus (Hebrew). We verify the quality of the UD annotation using an inter-annotator agreement study, and manually evaluate the transduced meaning representations. We then demonstrate the utility of the compiled corpora through (1) a longitudinal corpus study of the prevalence of different syntactic and semantic phenomena in the CDS, and (2) applying an existing computational model of language acquisition to the two corpora and briefly comparing the results across languages.
format	Preprint
id	arxiv_https___arxiv_org_abs_2109_10952
institution	arXiv
publishDate	2021
record_format	arxiv
spellingShingle	Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech Szubert, Ida Abend, Omri Schneider, Nathan Gibbon, Samuel Mahon, Louis Goldwater, Sharon Steedman, Mark Computation and Language This paper proposes a methodology for constructing such corpora of child directed speech (CDS) paired with sentential logical forms, and uses this method to create two such corpora, in English and Hebrew. The approach enforces a cross-linguistically consistent representation, building on recent advances in dependency representation and semantic parsing. Specifically, the approach involves two steps. First, we annotate the corpora using the Universal Dependencies (UD) scheme for syntactic annotation, which has been developed to apply consistently to a wide variety of domains and typologically diverse languages. Next, we further annotate these data by applying an automatic method for transducing sentential logical forms (LFs) from UD structures. The UD and LF representations have complementary strengths: UD structures are language-neutral and support consistent and reliable annotation by multiple annotators, whereas LFs are neutral as to their syntactic derivation and transparently encode semantic relations. Using this approach, we provide syntactic and semantic annotation for two corpora from CHILDES: Brown's Adam corpus (English; we annotate ~80% of its child-directed utterances), all child-directed utterances from Berman's Hagar corpus (Hebrew). We verify the quality of the UD annotation using an inter-annotator agreement study, and manually evaluate the transduced meaning representations. We then demonstrate the utility of the compiled corpora through (1) a longitudinal corpus study of the prevalence of different syntactic and semantic phenomena in the CDS, and (2) applying an existing computational model of language acquisition to the two corpora and briefly comparing the results across languages.
title	Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech
topic	Computation and Language
url	https://arxiv.org/abs/2109.10952

Ähnliche Einträge