Rapid Word Learning Through Meta In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Wentao, Jiang, Guangyuan, Linzen, Tal, Lake, Brenden M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912569828048896
author Wang, Wentao
Jiang, Guangyuan
Linzen, Tal
Lake, Brenden M.
author_facet Wang, Wentao
Jiang, Guangyuan
Linzen, Tal
Lake, Brenden M.
contents Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts. Yet the abilities of current language models for few-shot word learning, and methods for improving these abilities, are underexplored. In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow). This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word. This training is repeated on many new words to develop a general word-learning ability. We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pre-trained on orders of magnitude more data. Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples. These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2502_14791
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Rapid Word Learning Through Meta In-Context Learning
Wang, Wentao
Jiang, Guangyuan
Linzen, Tal
Lake, Brenden M.
Computation and Language
Artificial Intelligence
Machine Learning
Humans can quickly learn a new word from a few illustrative examples, and then systematically and flexibly use it in novel contexts. Yet the abilities of current language models for few-shot word learning, and methods for improving these abilities, are underexplored. In this study, we introduce a novel method, Meta-training for IN-context learNing Of Words (Minnow). This method trains language models to generate new examples of a word's usage given a few in-context examples, using a special placeholder token to represent the new word. This training is repeated on many new words to develop a general word-learning ability. We find that training models from scratch with Minnow on human-scale child-directed language enables strong few-shot word learning, comparable to a large language model (LLM) pre-trained on orders of magnitude more data. Furthermore, through discriminative and generative evaluations, we demonstrate that finetuning pre-trained LLMs with Minnow improves their ability to discriminate between new words, identify syntactic categories of new words, and generate reasonable new usages and definitions for new words, based on one or a few in-context examples. These findings highlight the data efficiency of Minnow and its potential to improve language model performance in word learning tasks.
title Rapid Word Learning Through Meta In-Context Learning
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2502.14791