Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Weissweiler, Leonie, Mahowald, Kyle, Goldberg, Adele
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909795489939456
author Weissweiler, Leonie
Mahowald, Kyle
Goldberg, Adele
author_facet Weissweiler, Leonie
Mahowald, Kyle
Goldberg, Adele
contents Linguistic evaluations of how well LMs generalize to produce or understand language often implicitly take for granted that natural languages are generated by symbolic rules. According to this perspective, grammaticality is determined by whether sentences obey such rules. Interpretation is compositionally generated by syntactic rules operating on meaningful words. Semantic parsing maps sentences into formal logic. Failures of LMs to obey strict rules are presumed to reveal that LMs do not produce or understand language like humans. Here we suggest that LMs' failures to obey symbolic rules may be a feature rather than a bug, because natural languages are not based on neatly separable, compositional rules. Rather, new utterances are produced and understood by a combination of flexible, interrelated, and context-dependent constructions. Considering gradient factors such as frequencies, context, and function will help us reimagine new benchmarks and analyses to probe whether and how LMs capture the rich, flexible generalizations that comprise natural languages.
format Preprint
id arxiv_https___arxiv_org_abs_2502_13195
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
Weissweiler, Leonie
Mahowald, Kyle
Goldberg, Adele
Computation and Language
Linguistic evaluations of how well LMs generalize to produce or understand language often implicitly take for granted that natural languages are generated by symbolic rules. According to this perspective, grammaticality is determined by whether sentences obey such rules. Interpretation is compositionally generated by syntactic rules operating on meaningful words. Semantic parsing maps sentences into formal logic. Failures of LMs to obey strict rules are presumed to reveal that LMs do not produce or understand language like humans. Here we suggest that LMs' failures to obey symbolic rules may be a feature rather than a bug, because natural languages are not based on neatly separable, compositional rules. Rather, new utterances are produced and understood by a combination of flexible, interrelated, and context-dependent constructions. Considering gradient factors such as frequencies, context, and function will help us reimagine new benchmarks and analyses to probe whether and how LMs capture the rich, flexible generalizations that comprise natural languages.
title Linguistic Generalizations are not Rules: Impacts on Evaluation of LMs
topic Computation and Language
url https://arxiv.org/abs/2502.13195