Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Jiuqi, Blaser, Ethan, Daneshmand, Hadi, Zhang, Shangtong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910842938720256
author Wang, Jiuqi
Blaser, Ethan
Daneshmand, Hadi
Zhang, Shangtong
author_facet Wang, Jiuqi
Blaser, Ethan
Daneshmand, Hadi
Zhang, Shangtong
contents Traditionally, reinforcement learning (RL) agents learn to solve new tasks by updating their neural network parameters through interactions with the task environment. However, recent works demonstrate that some RL agents, after certain pretraining procedures, can learn to solve unseen new tasks without parameter updates, a phenomenon known as in-context reinforcement learning (ICRL). The empirical success of ICRL is widely attributed to the hypothesis that the forward pass of the pretrained agent neural network implements an RL algorithm. In this paper, we support this hypothesis by showing, both empirically and theoretically, that when a transformer is trained for policy evaluation tasks, it can discover and learn to implement temporal difference learning in its forward pass.
format Preprint
id arxiv_https___arxiv_org_abs_2405_13861
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
Wang, Jiuqi
Blaser, Ethan
Daneshmand, Hadi
Zhang, Shangtong
Machine Learning
Traditionally, reinforcement learning (RL) agents learn to solve new tasks by updating their neural network parameters through interactions with the task environment. However, recent works demonstrate that some RL agents, after certain pretraining procedures, can learn to solve unseen new tasks without parameter updates, a phenomenon known as in-context reinforcement learning (ICRL). The empirical success of ICRL is widely attributed to the hypothesis that the forward pass of the pretrained agent neural network implements an RL algorithm. In this paper, we support this hypothesis by showing, both empirically and theoretically, that when a transformer is trained for policy evaluation tasks, it can discover and learn to implement temporal difference learning in its forward pass.
title Transformers Can Learn Temporal Difference Methods for In-Context Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2405.13861