Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Misra, Kanishka, Mahowald, Kyle
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912447987712000
author Misra, Kanishka
Mahowald, Kyle
author_facet Misra, Kanishka
Mahowald, Kyle
contents Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question. To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction (``a beautiful five days''). We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed. We found that AANNs were still learned better than systematically perturbed variants of the construction. Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., ``a few days''). An additional experiment showed that this learning is enhanced when there is more variability in the input. Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena. Data and code: https://github.com/kanishkamisra/aannalysis.
format Preprint
id arxiv_https___arxiv_org_abs_2403_19827
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
Misra, Kanishka
Mahowald, Kyle
Computation and Language
Language models learn rare syntactic phenomena, but the extent to which this is attributable to generalization vs. memorization is a major open question. To that end, we iteratively trained transformer language models on systematically manipulated corpora which were human-scale in size, and then evaluated their learning of a rare grammatical phenomenon: the English Article+Adjective+Numeral+Noun (AANN) construction (``a beautiful five days''). We compared how well this construction was learned on the default corpus relative to a counterfactual corpus in which AANN sentences were removed. We found that AANNs were still learned better than systematically perturbed variants of the construction. Using additional counterfactual corpora, we suggest that this learning occurs through generalization from related constructions (e.g., ``a few days''). An additional experiment showed that this learning is enhanced when there is more variability in the input. Taken together, our results provide an existence proof that LMs can learn rare grammatical phenomena by generalization from less rare phenomena. Data and code: https://github.com/kanishkamisra/aannalysis.
title Language Models Learn Rare Phenomena from Less Rare Phenomena: The Case of the Missing AANNs
topic Computation and Language
url https://arxiv.org/abs/2403.19827