Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hsu, Alexander, Shen, Zhaiming, Liao, Wenjing, Lai, Rongjie
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915984589193216
author Hsu, Alexander
Shen, Zhaiming
Liao, Wenjing
Lai, Rongjie
author_facet Hsu, Alexander
Shen, Zhaiming
Liao, Wenjing
Lai, Rongjie
contents Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2605_05176
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
Hsu, Alexander
Shen, Zhaiming
Liao, Wenjing
Lai, Rongjie
Machine Learning
Numerical Analysis
Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
title Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer
topic Machine Learning
Numerical Analysis
url https://arxiv.org/abs/2605.05176