In-Context Deep Learning via Transformer Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Weimin, Su, Maojiang, Hu, Jerry Yao-Chieh, Song, Zhao, Liu, Han
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913789076570112
author Wu, Weimin
Su, Maojiang
Hu, Jerry Yao-Chieh
Song, Zhao
Liu, Han
author_facet Wu, Weimin
Su, Maojiang
Hu, Jerry Yao-Chieh
Song, Zhao
Liu, Han
contents We investigate the transformer's capability to simulate the training process of deep models via in-context learning (ICL), i.e., in-context deep learning. Our key contribution is providing a positive example of using a transformer to train a deep neural network by gradient descent in an implicit fashion via ICL. Specifically, we provide an explicit construction of a $(2N+4)L$-layer transformer capable of simulating $L$ gradient descent steps of an $N$-layer ReLU network through ICL. We also give the theoretical guarantees for the approximation within any given error and the convergence of the ICL gradient descent. Additionally, we extend our analysis to the more practical setting using Softmax-based transformers. We validate our findings on synthetic datasets for 3-layer, 4-layer, and 6-layer neural networks. The results show that ICL performance matches that of direct training.
format Preprint
id arxiv_https___arxiv_org_abs_2411_16549
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle In-Context Deep Learning via Transformer Models
Wu, Weimin
Su, Maojiang
Hu, Jerry Yao-Chieh
Song, Zhao
Liu, Han
Machine Learning
We investigate the transformer's capability to simulate the training process of deep models via in-context learning (ICL), i.e., in-context deep learning. Our key contribution is providing a positive example of using a transformer to train a deep neural network by gradient descent in an implicit fashion via ICL. Specifically, we provide an explicit construction of a $(2N+4)L$-layer transformer capable of simulating $L$ gradient descent steps of an $N$-layer ReLU network through ICL. We also give the theoretical guarantees for the approximation within any given error and the convergence of the ICL gradient descent. Additionally, we extend our analysis to the more practical setting using Softmax-based transformers. We validate our findings on synthetic datasets for 3-layer, 4-layer, and 6-layer neural networks. The results show that ICL performance matches that of direct training.
title In-Context Deep Learning via Transformer Models
topic Machine Learning
url https://arxiv.org/abs/2411.16549