For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Cuneo, Nicole, Graves, Eleanor, Rakshit, Supantho, Goldberg, Adele E.
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913835339743232
author Cuneo, Nicole
Graves, Eleanor
Rakshit, Supantho
Goldberg, Adele E.
author_facet Cuneo, Nicole
Graves, Eleanor
Rakshit, Supantho
Goldberg, Adele E.
contents It remains debated how well any LM understands natural language or generates reliable metalinguistic judgments. Moreover, relatively little work has demonstrated that LMs can represent and respect subtle relationships between form and function proposed by linguists. We here focus on a particular such relationship established in recent work: English speakers' judgments about the information structure of canonical sentences predicts independently collected acceptability ratings on corresponding 'long distance dependency' [LDD] constructions, across a wide array of base constructions and multiple types of LDDs. To determine whether any LM captures this relationship, we probe GPT-4 on the same tasks used with humans and new extensions.Results reveal reliable metalinguistic skill on the information structure and acceptability tasks, replicating a striking interaction between the two, despite the zero-shot, explicit nature of the tasks, and little to no chance of contamination [Studies 1a, 1b]. Study 2 manipulates the information structure of base sentences and confirms a causal relationship: increasing the prominence of a constituent in a context sentence increases the subsequent acceptability ratings on an LDD construction. The findings suggest a tight relationship between natural and GPT-4 generated English, and between information structure and syntax, which begs for further exploration.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09005
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies
Cuneo, Nicole
Graves, Eleanor
Rakshit, Supantho
Goldberg, Adele E.
Computation and Language
It remains debated how well any LM understands natural language or generates reliable metalinguistic judgments. Moreover, relatively little work has demonstrated that LMs can represent and respect subtle relationships between form and function proposed by linguists. We here focus on a particular such relationship established in recent work: English speakers' judgments about the information structure of canonical sentences predicts independently collected acceptability ratings on corresponding 'long distance dependency' [LDD] constructions, across a wide array of base constructions and multiple types of LDDs. To determine whether any LM captures this relationship, we probe GPT-4 on the same tasks used with humans and new extensions.Results reveal reliable metalinguistic skill on the information structure and acceptability tasks, replicating a striking interaction between the two, despite the zero-shot, explicit nature of the tasks, and little to no chance of contamination [Studies 1a, 1b]. Study 2 manipulates the information structure of base sentences and confirms a causal relationship: increasing the prominence of a constituent in a context sentence increases the subsequent acceptability ratings on an LDD construction. The findings suggest a tight relationship between natural and GPT-4 generated English, and between information structure and syntax, which begs for further exploration.
title For GPT-4 as with Humans: Information Structure Predicts Acceptability of Long-Distance Dependencies
topic Computation and Language
url https://arxiv.org/abs/2505.09005