Saved in:
Bibliographic Details
Main Author: Johnsen, Lars G. B.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2512.10453
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911313046798336
author Johnsen, Lars G. B.
author_facet Johnsen, Lars G. B.
contents What counts as evidence for syntactic structure? In traditional generative grammar, systematic contrasts in grammaticality such as subject-auxiliary inversion and the licensing of parasitic gaps are taken as evidence for an internal, hierarchical grammar. In this paper, we test whether large language models (LLMs), trained only on surface forms, reproduce these contrasts in ways that imply an underlying structural representation. We focus on two classic constructions: subject-auxiliary inversion (testing recognition of the subject boundary) and parasitic gap licensing (testing abstract dependency structure). We evaluate models including GPT-4 and LLaMA-3 using prompts eliciting acceptability ratings. Results show that LLMs reliably distinguish between grammatical and ungrammatical variants in both constructions, and as such support that they are sensitive to structure and not just linear order. Structural generalizations, distinct from cognitive knowledge, emerge from predictive training on surface forms, suggesting functional sensitivity to syntax without explicit encoding.
format Preprint
id arxiv_https___arxiv_org_abs_2512_10453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs
Johnsen, Lars G. B.
Computation and Language
What counts as evidence for syntactic structure? In traditional generative grammar, systematic contrasts in grammaticality such as subject-auxiliary inversion and the licensing of parasitic gaps are taken as evidence for an internal, hierarchical grammar. In this paper, we test whether large language models (LLMs), trained only on surface forms, reproduce these contrasts in ways that imply an underlying structural representation. We focus on two classic constructions: subject-auxiliary inversion (testing recognition of the subject boundary) and parasitic gap licensing (testing abstract dependency structure). We evaluate models including GPT-4 and LLaMA-3 using prompts eliciting acceptability ratings. Results show that LLMs reliably distinguish between grammatical and ungrammatical variants in both constructions, and as such support that they are sensitive to structure and not just linear order. Structural generalizations, distinct from cognitive knowledge, emerge from predictive training on surface forms, suggesting functional sensitivity to syntax without explicit encoding.
title Grammaticality Judgments in Humans and Language Models: Revisiting Generative Grammar with LLMs
topic Computation and Language
url https://arxiv.org/abs/2512.10453