Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Baiyuan, Ito, Shinji, Imaizumi, Masaaki
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918166270050304
author Chen, Baiyuan
Ito, Shinji
Imaizumi, Masaaki
author_facet Chen, Baiyuan
Ito, Shinji
Imaizumi, Masaaki
contents Transformers have demonstrated exceptional performance across a wide range of domains. While their ability to perform reinforcement learning in-context has been established both theoretically and empirically, their behavior in non-stationary environments remains less understood. In this study, we address this gap by showing that transformers can achieve nearly optimal dynamic regret bounds in non-stationary settings. We prove that transformers are capable of approximating strategies used to handle non-stationary environments and can learn the approximator in the in-context learning setup. Our experiments further show that transformers can match or even outperform existing expert algorithms in such environments.
format Preprint
id arxiv_https___arxiv_org_abs_2508_16027
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
Chen, Baiyuan
Ito, Shinji
Imaizumi, Masaaki
Machine Learning
Transformers have demonstrated exceptional performance across a wide range of domains. While their ability to perform reinforcement learning in-context has been established both theoretically and empirically, their behavior in non-stationary environments remains less understood. In this study, we address this gap by showing that transformers can achieve nearly optimal dynamic regret bounds in non-stationary settings. We prove that transformers are capable of approximating strategies used to handle non-stationary environments and can learn the approximator in the in-context learning setup. Our experiments further show that transformers can match or even outperform existing expert algorithms in such environments.
title Optimal Dynamic Regret by Transformers for Non-Stationary Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2508.16027