Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Kobayashi, Seijin, Schimpf, Yanick, Schlegel, Maximilian, Steger, Angelika, Wolczyk, Maciej, von Oswald, Johannes, Scherrer, Nino, Maile, Kaitlin, Lajoie, Guillaume, Richards, Blake A., Saurous, Rif A., Manyika, James, Arcas, Blaise Agüera y, Meulemans, Alexander, Sacramento, João
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866917168503848960
author Kobayashi, Seijin
Schimpf, Yanick
Schlegel, Maximilian
Steger, Angelika
Wolczyk, Maciej
von Oswald, Johannes
Scherrer, Nino
Maile, Kaitlin
Lajoie, Guillaume
Richards, Blake A.
Saurous, Rif A.
Manyika, James
Arcas, Blaise Agüera y
Meulemans, Alexander
Sacramento, João
author_facet Kobayashi, Seijin
Schimpf, Yanick
Schlegel, Maximilian
Steger, Angelika
Wolczyk, Maciej
von Oswald, Johannes
Scherrer, Nino
Maile, Kaitlin
Lajoie, Guillaume
Richards, Blake A.
Saurous, Rif A.
Manyika, James
Arcas, Blaise Agüera y
Meulemans, Alexander
Sacramento, João
contents Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains. During RL, these models explore by generating new outputs, one token at a time. However, sampling actions token-by-token can result in highly inefficient learning, particularly when rewards are sparse. Here, we show that it is possible to overcome this problem by acting and exploring within the internal representations of an autoregressive model. Specifically, to discover temporally-abstract actions, we introduce a higher-order, non-causal sequence model whose outputs control the residual stream activations of a base autoregressive model. On grid world and MuJoCo-based tasks with hierarchical structure, we find that the higher-order model learns to compress long activation sequence chunks onto internal controllers. Critically, each controller executes a sequence of behaviorally meaningful actions that unfold over long timescales and are accompanied with a learned termination condition, such that composing multiple controllers over time leads to efficient exploration on novel tasks. We show that direct internal controller reinforcement, a process we term "internal RL", enables learning from sparse rewards in cases where standard RL finetuning fails. Our results demonstrate the benefits of latent action generation and reinforcement in autoregressive models, suggesting internal RL as a promising avenue for realizing hierarchical RL within foundation models.
format Preprint
id arxiv_https___arxiv_org_abs_2512_20605
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
Kobayashi, Seijin
Schimpf, Yanick
Schlegel, Maximilian
Steger, Angelika
Wolczyk, Maciej
von Oswald, Johannes
Scherrer, Nino
Maile, Kaitlin
Lajoie, Guillaume
Richards, Blake A.
Saurous, Rif A.
Manyika, James
Arcas, Blaise Agüera y
Meulemans, Alexander
Sacramento, João
Machine Learning
Artificial Intelligence
Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains. During RL, these models explore by generating new outputs, one token at a time. However, sampling actions token-by-token can result in highly inefficient learning, particularly when rewards are sparse. Here, we show that it is possible to overcome this problem by acting and exploring within the internal representations of an autoregressive model. Specifically, to discover temporally-abstract actions, we introduce a higher-order, non-causal sequence model whose outputs control the residual stream activations of a base autoregressive model. On grid world and MuJoCo-based tasks with hierarchical structure, we find that the higher-order model learns to compress long activation sequence chunks onto internal controllers. Critically, each controller executes a sequence of behaviorally meaningful actions that unfold over long timescales and are accompanied with a learned termination condition, such that composing multiple controllers over time leads to efficient exploration on novel tasks. We show that direct internal controller reinforcement, a process we term "internal RL", enables learning from sparse rewards in cases where standard RL finetuning fails. Our results demonstrate the benefits of latent action generation and reinforcement in autoregressive models, suggesting internal RL as a promising avenue for realizing hierarchical RL within foundation models.
title Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2512.20605