Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Taufeeque, Mohammad, Tucker, Aaron David, Gleave, Adam, Garriga-Alonso, Adrià
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913164817334272
author Taufeeque, Mohammad
Tucker, Aaron David
Gleave, Adam
Garriga-Alonso, Adrià
author_facet Taufeeque, Mohammad
Tucker, Aaron David
Gleave, Adam
Garriga-Alonso, Adrià
contents We partially reverse-engineer a convolutional recurrent neural network (RNN) trained with model-free reinforcement learning to play the box-pushing game Sokoban. We find that the RNN stores future moves (plans) as activations in particular channels of the hidden state, which we call path channels. A high activation in a particular location means that, when a box is in that location, it will get pushed in the channel's assigned direction. We examine the convolutional kernels between path channels and find that they encode the change in position resulting from each possible action, thus representing part of a learned transition model. The RNN constructs plans by starting at the boxes and goals. These kernels extend activations in path channels forwards from boxes and backwards from the goal. Negative values are placed in channels at obstacles. This causes the extension kernels to propagate the negative value in reverse, thus pruning the last few steps and letting an alternative plan emerge; a form of backtracking. Our work shows that, a precise understanding of the plan representation allows us to directly understand the bidirectional planning-like algorithm learned by model-free training in more familiar terms.
format Preprint
id arxiv_https___arxiv_org_abs_2506_10138
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
Taufeeque, Mohammad
Tucker, Aaron David
Gleave, Adam
Garriga-Alonso, Adrià
Machine Learning
Artificial Intelligence
We partially reverse-engineer a convolutional recurrent neural network (RNN) trained with model-free reinforcement learning to play the box-pushing game Sokoban. We find that the RNN stores future moves (plans) as activations in particular channels of the hidden state, which we call path channels. A high activation in a particular location means that, when a box is in that location, it will get pushed in the channel's assigned direction. We examine the convolutional kernels between path channels and find that they encode the change in position resulting from each possible action, thus representing part of a learned transition model. The RNN constructs plans by starting at the boxes and goals. These kernels extend activations in path channels forwards from boxes and backwards from the goal. Negative values are placed in channels at obstacles. This causes the extension kernels to propagate the negative value in reverse, thus pruning the last few steps and letting an alternative plan emerge; a form of backtracking. Our work shows that, a precise understanding of the plan representation allows us to directly understand the bidirectional planning-like algorithm learned by model-free training in more familiar terms.
title Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.10138