Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xingxuan, Wang, Haoran, Li, Jiansheng, Xue, Yuan, Guan, Shikai, Xu, Renzhe, Zou, Hao, Yu, Han, Cui, Peng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912283837333504
author Zhang, Xingxuan
Wang, Haoran
Li, Jiansheng
Xue, Yuan
Guan, Shikai
Xu, Renzhe
Zou, Hao
Yu, Han
Cui, Peng
author_facet Zhang, Xingxuan
Wang, Haoran
Li, Jiansheng
Xue, Yuan
Guan, Shikai
Xu, Renzhe
Zou, Hao
Yu, Han
Cui, Peng
contents Large language models (LLMs) like GPT-4 and LLaMA-3 utilize the powerful in-context learning (ICL) capability of Transformer architecture to learn on the fly from limited examples. While ICL underpins many LLM applications, its full potential remains hindered by a limited understanding of its generalization boundaries and vulnerabilities. We present a systematic investigation of transformers' generalization capability with ICL relative to training data coverage by defining a task-centric framework along three dimensions: inter-problem, intra-problem, and intra-task generalization. Through extensive simulation and real-world experiments, encompassing tasks such as function fitting, API calling, and translation, we find that transformers lack inter-problem generalization with ICL, but excel in intra-task and intra-problem generalization. When the training data includes a greater variety of mixed tasks, it significantly enhances the generalization ability of ICL on unseen tasks and even on known simple tasks. This guides us in designing training data to maximize the diversity of tasks covered and to combine different tasks whenever possible, rather than solely focusing on the target task for testing.
format Preprint
id arxiv_https___arxiv_org_abs_2503_15579
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
Zhang, Xingxuan
Wang, Haoran
Li, Jiansheng
Xue, Yuan
Guan, Shikai
Xu, Renzhe
Zou, Hao
Yu, Han
Cui, Peng
Machine Learning
Large language models (LLMs) like GPT-4 and LLaMA-3 utilize the powerful in-context learning (ICL) capability of Transformer architecture to learn on the fly from limited examples. While ICL underpins many LLM applications, its full potential remains hindered by a limited understanding of its generalization boundaries and vulnerabilities. We present a systematic investigation of transformers' generalization capability with ICL relative to training data coverage by defining a task-centric framework along three dimensions: inter-problem, intra-problem, and intra-task generalization. Through extensive simulation and real-world experiments, encompassing tasks such as function fitting, API calling, and translation, we find that transformers lack inter-problem generalization with ICL, but excel in intra-task and intra-problem generalization. When the training data includes a greater variety of mixed tasks, it significantly enhances the generalization ability of ICL on unseen tasks and even on known simple tasks. This guides us in designing training data to maximize the diversity of tasks covered and to combine different tasks whenever possible, rather than solely focusing on the target task for testing.
title Understanding the Generalization of In-Context Learning in Transformers: An Empirical Study
topic Machine Learning
url https://arxiv.org/abs/2503.15579