An Information-Theoretic Analysis of In-Context Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Jeon, Hong Jun, Lee, Jason D., Lei, Qi, Van Roy, Benjamin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913213074898944
author Jeon, Hong Jun
Lee, Jason D.
Lei, Qi
Van Roy, Benjamin
author_facet Jeon, Hong Jun
Lee, Jason D.
Lei, Qi
Van Roy, Benjamin
contents Previous theoretical results pertaining to meta-learning on sequences build on contrived assumptions and are somewhat convoluted. We introduce new information-theoretic tools that lead to an elegant and very general decomposition of error into three components: irreducible error, meta-learning error, and intra-task error. These tools unify analyses across many meta-learning challenges. To illustrate, we apply them to establish new results about in-context learning with transformers. Our theoretical results characterizes how error decays in both the number of training sequences and sequence lengths. Our results are very general; for example, they avoid contrived mixing time assumptions made by all prior results that establish decay of error with sequence length.
format Preprint
id arxiv_https___arxiv_org_abs_2401_15530
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle An Information-Theoretic Analysis of In-Context Learning
Jeon, Hong Jun
Lee, Jason D.
Lei, Qi
Van Roy, Benjamin
Machine Learning
Information Theory
Previous theoretical results pertaining to meta-learning on sequences build on contrived assumptions and are somewhat convoluted. We introduce new information-theoretic tools that lead to an elegant and very general decomposition of error into three components: irreducible error, meta-learning error, and intra-task error. These tools unify analyses across many meta-learning challenges. To illustrate, we apply them to establish new results about in-context learning with transformers. Our theoretical results characterizes how error decays in both the number of training sequences and sequence lengths. Our results are very general; for example, they avoid contrived mixing time assumptions made by all prior results that establish decay of error with sequence length.
title An Information-Theoretic Analysis of In-Context Learning
topic Machine Learning
Information Theory
url https://arxiv.org/abs/2401.15530