Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhao, Yike, Eberhard, Onno, Khammassi, Malek, Sayed, Ali H., Muehlebach, Michael
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917547851382784
author Zhao, Yike
Eberhard, Onno
Khammassi, Malek
Sayed, Ali H.
Muehlebach, Michael
author_facet Zhao, Yike
Eberhard, Onno
Khammassi, Malek
Sayed, Ali H.
Muehlebach, Michael
contents The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning. We provide a theoretical justification for their empirical effectiveness by constructing and studying two linear filters: (i) the first exactly reproduces the pre-softmax logits of the belief vector in a hidden Markov model (HMM) under a deterministic transition matrix, thereby serving as a sufficient statistic for optimal policy learning, (ii) the second achieves vanishing state-decoding error under a nearly deterministic transition matrix, thus reducing state ambiguity to near zero. The results extend to action-controlled HMMs, where the corresponding linear filters become time-varying with action-dependent dynamics. We illustrate our main results through numerical experiments and further show that the constructed linear filter serves as a strong feature extractor in a small reinforcement learning game.
format Preprint
id arxiv_https___arxiv_org_abs_2605_31261
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
Zhao, Yike
Eberhard, Onno
Khammassi, Malek
Sayed, Ali H.
Muehlebach, Michael
Machine Learning
Artificial Intelligence
The family of linear recurrent neural networks has shown strong performance as recurrent memory units in partially observable reinforcement learning. We provide a theoretical justification for their empirical effectiveness by constructing and studying two linear filters: (i) the first exactly reproduces the pre-softmax logits of the belief vector in a hidden Markov model (HMM) under a deterministic transition matrix, thereby serving as a sufficient statistic for optimal policy learning, (ii) the second achieves vanishing state-decoding error under a nearly deterministic transition matrix, thus reducing state ambiguity to near zero. The results extend to action-controlled HMMs, where the corresponding linear filters become time-varying with action-dependent dynamics. We illustrate our main results through numerical experiments and further show that the constructed linear filter serves as a strong feature extractor in a small reinforcement learning game.
title Why Linear Recurrent Memory Works in Partially Observable Reinforcement Learning
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.31261