Re-examining learning linear functions in context

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Naim, Omar, Fouilhé, Guilhem, Asher, Nicholas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911130734034944
author Naim, Omar
Fouilhé, Guilhem
Asher, Nicholas
author_facet Naim, Omar
Fouilhé, Guilhem
Asher, Nicholas
contents In-context learning (ICL) has emerged as a powerful paradigm for easily adapting Large Language Models (LLMs) to various tasks. However, our understanding of how ICL works remains limited. We explore a simple model of ICL in a controlled setup with synthetic training data to investigate ICL of univariate linear functions. We experiment with a range of GPT-2-like transformer models trained from scratch. Our findings challenge the prevailing narrative that transformers adopt algorithmic approaches like linear regression to learn a linear function in-context. These models fail to generalize beyond their training distribution, highlighting fundamental limitations in their capacity to infer abstract task structures. Our experiments lead us to propose a mathematically precise hypothesis of what the model might be learning.
format Preprint
id arxiv_https___arxiv_org_abs_2411_11465
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Re-examining learning linear functions in context
Naim, Omar
Fouilhé, Guilhem
Asher, Nicholas
Machine Learning
Computation and Language
In-context learning (ICL) has emerged as a powerful paradigm for easily adapting Large Language Models (LLMs) to various tasks. However, our understanding of how ICL works remains limited. We explore a simple model of ICL in a controlled setup with synthetic training data to investigate ICL of univariate linear functions. We experiment with a range of GPT-2-like transformer models trained from scratch. Our findings challenge the prevailing narrative that transformers adopt algorithmic approaches like linear regression to learn a linear function in-context. These models fail to generalize beyond their training distribution, highlighting fundamental limitations in their capacity to infer abstract task structures. Our experiments lead us to propose a mathematically precise hypothesis of what the model might be learning.
title Re-examining learning linear functions in context
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2411.11465