HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dang, Long H, Rawlinson, David
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911234155085824
author Dang, Long H
Rawlinson, David
author_facet Dang, Long H
Rawlinson, David
contents The Hierarchical Reasoning Model (HRM) has impressive reasoning abilities given its small size, but has only been applied to supervised, static, fully-observable problems. One of HRM's strengths is its ability to adapt its computational effort to the difficulty of the problem. However, in its current form it cannot integrate and reuse computation from previous time-steps if the problem is dynamic, uncertain or partially observable, or be applied where the correct action is undefined, characteristics of many real-world problems. This paper presents HRM-Agent, a variant of HRM trained using only reinforcement learning. We show that HRM can learn to navigate to goals in dynamic and uncertain maze environments. Recent work suggests that HRM's reasoning abilities stem from its recurrent inference process. We explore the dynamics of the recurrent inference process and find evidence that it is successfully reusing computation from earlier environment time-steps.
format Preprint
id arxiv_https___arxiv_org_abs_2510_22832
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
Dang, Long H
Rawlinson, David
Artificial Intelligence
Machine Learning
68T07 (Primary) 62M45, 37N99 (Secondary)
I.2.6; I.2.8
The Hierarchical Reasoning Model (HRM) has impressive reasoning abilities given its small size, but has only been applied to supervised, static, fully-observable problems. One of HRM's strengths is its ability to adapt its computational effort to the difficulty of the problem. However, in its current form it cannot integrate and reuse computation from previous time-steps if the problem is dynamic, uncertain or partially observable, or be applied where the correct action is undefined, characteristics of many real-world problems. This paper presents HRM-Agent, a variant of HRM trained using only reinforcement learning. We show that HRM can learn to navigate to goals in dynamic and uncertain maze environments. Recent work suggests that HRM's reasoning abilities stem from its recurrent inference process. We explore the dynamics of the recurrent inference process and find evidence that it is successfully reusing computation from earlier environment time-steps.
title HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
topic Artificial Intelligence
Machine Learning
68T07 (Primary) 62M45, 37N99 (Secondary)
I.2.6; I.2.8
url https://arxiv.org/abs/2510.22832