HRM-Agent: Training a recurrent reasoning model in dynamic environments using reinforcement learning
Fuente:
arXiv
Saved in:
| Main Authors: | Dang, Long H, Rawlinson, David |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026)
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
by: Caunhye, Ali Murtaza, et al.
Published: (2025)
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026)
by: Duy, Dang Sy, et al.
Published: (2026)
Soil Compaction Parameters Prediction Based on Automated Machine Learning Approach
by: Erden, Caner, et al.
Published: (2025)
by: Erden, Caner, et al.
Published: (2025)
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)
by: Saini, Saurabh, et al.
Published: (2026)
UWM-JEPA: Predictive World Models That Imagine in Belief Space
by: Radha, Santosh Kumar, et al.
Published: (2026)
by: Radha, Santosh Kumar, et al.
Published: (2026)
TACIT: Transformation-Aware Capturing of Implicit Thought
by: Nobrega, Daniel
Published: (2026)
by: Nobrega, Daniel
Published: (2026)
Emotion-Inspired Learning Signals (EILS): A Homeostatic Framework for Adaptive Autonomous Agents
by: Tiwari, Dhruv
Published: (2025)
by: Tiwari, Dhruv
Published: (2025)
From Imitation to Interaction: Mastering Game of Schnapsen with Shallow Reinforcement Learning
by: Klačan, Ján, et al.
Published: (2026)
by: Klačan, Ján, et al.
Published: (2026)
AgentDS Technical Report: Benchmarking the Future of Human-AI Collaboration in Domain-Specific Data Science
by: Luo, An, et al.
Published: (2026)
by: Luo, An, et al.
Published: (2026)
Multimodal Generative AI for Story Point Estimation in Software Development
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
by: Islam, Mohammad Rubyet, et al.
Published: (2025)
Improving Efficiency of Sampling-based Motion Planning via Message-Passing Monte Carlo
by: Chahine, Makram, et al.
Published: (2024)
by: Chahine, Makram, et al.
Published: (2024)
CellARC: Measuring Intelligence with Cellular Automata
by: Lžičař, Miroslav
Published: (2025)
by: Lžičař, Miroslav
Published: (2025)
Ice Cream Doesn't Cause Drowning: Benchmarking LLMs Against Statistical Pitfalls in Causal Inference
by: Du, Jin, et al.
Published: (2025)
by: Du, Jin, et al.
Published: (2025)
Autoencoders in Function Space
by: Bunker, Justin, et al.
Published: (2024)
by: Bunker, Justin, et al.
Published: (2024)
Context Engineering for Multi-Agent LLM Code Assistants Using Elicit, NotebookLM, ChatGPT, and Claude Code
by: Haseeb, Muhammad
Published: (2025)
by: Haseeb, Muhammad
Published: (2025)
KL-Regularised Q-Learning: A Token-level Action-Value perspective on Online RLHF
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
Observable Neural ODEs for Identifiable Causal Forecasting in Continuous Time
by: Wendland, Jennifer, et al.
Published: (2026)
by: Wendland, Jennifer, et al.
Published: (2026)
Learning from Preferences and Mixed Demonstrations in General Settings
by: Brown, Jason R, et al.
Published: (2025)
by: Brown, Jason R, et al.
Published: (2025)
LLM-Assisted Iterative Evolution with Swarm Intelligence Toward SuperBrain
by: Weigang, Li, et al.
Published: (2025)
by: Weigang, Li, et al.
Published: (2025)
Can Agentic AI Match the Performance of Human Data Scientists?
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
Self-Supervised On-Policy Reinforcement Learning via Contrastive Proximal Policy Optimisation
by: Osman, Asim, et al.
Published: (2026)
by: Osman, Asim, et al.
Published: (2026)
Generalization and Feature Attribution in Machine Learning Models for Crop Yield and Anomaly Prediction in Germany
by: Baatz, Roland
Published: (2025)
by: Baatz, Roland
Published: (2025)
Adaptive Negative Scheduling for Graph Contrastive Learning
by: Ali, Adnan, et al.
Published: (2026)
by: Ali, Adnan, et al.
Published: (2026)
The Impact of Data Characteristics on GNN Evaluation for Detecting Fake News
by: Karn, Isha, et al.
Published: (2025)
by: Karn, Isha, et al.
Published: (2025)
Gaussian Ensemble Belief Propagation for Efficient Inference in High-Dimensional Systems
by: MacKinlay, Dan, et al.
Published: (2024)
by: MacKinlay, Dan, et al.
Published: (2024)
Memory-Augmented State Machine Prompting: A Novel LLM Agent Framework for Real-Time Strategy Games
by: Qi, Runnan, et al.
Published: (2025)
by: Qi, Runnan, et al.
Published: (2025)
Image-based Facial Rig Inversion
by: Yang, Tianxiang, et al.
Published: (2025)
by: Yang, Tianxiang, et al.
Published: (2025)
Can a Bayesian Oracle Prevent Harm from an Agent?
by: Bengio, Yoshua, et al.
Published: (2024)
by: Bengio, Yoshua, et al.
Published: (2024)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
iLTM: Integrated Large Tabular Model
by: Bonet, David, et al.
Published: (2025)
by: Bonet, David, et al.
Published: (2025)
AssistedDS: Benchmarking How External Domain Knowledge Assists LLMs in Automated Data Science
by: Luo, An, et al.
Published: (2025)
by: Luo, An, et al.
Published: (2025)
Soft Actor-Critic with Beta Policy via Implicit Reparameterization Gradients
by: Della Libera, Luca
Published: (2024)
by: Della Libera, Luca
Published: (2024)
Multiple data-driven missing imputation
by: Kavun, Sergii
Published: (2025)
by: Kavun, Sergii
Published: (2025)
On the Fundamental Limitations of Decentralized Learnable Reward Shaping in Cooperative Multi-Agent Reinforcement Learning
by: Akella, Aditya
Published: (2025)
by: Akella, Aditya
Published: (2025)
ElasticFlow: One-Step Physics-Consistent Policy with Elastic Time Horizons for Language-Guided Manipulation
by: Chen, Kewei, et al.
Published: (2026)
by: Chen, Kewei, et al.
Published: (2026)
CogniLoad: A Synthetic Natural Language Reasoning Benchmark With Tunable Length, Intrinsic Difficulty, and Distractor Density
by: Kaiser, Daniel, et al.
Published: (2025)
by: Kaiser, Daniel, et al.
Published: (2025)
The unknotting number, hard unknot diagrams, and reinforcement learning
by: Applebaum, Taylor, et al.
Published: (2024)
by: Applebaum, Taylor, et al.
Published: (2024)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Evolving machine learning workflows through interactive AutoML
by: Barbudo, Rafael, et al.
Published: (2024)
by: Barbudo, Rafael, et al.
Published: (2024)
Similar Items
-
Urban Spatio-Temporal Foundation Models for Climate-Resilient Housing: Scaling Diffusion Transformers for Disaster Risk Prediction
by: Imanov, Olaf Yunus Laitinen, et al.
Published: (2026) -
A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms
by: Caunhye, Ali Murtaza, et al.
Published: (2025) -
Normalisation and Initialisation Strategies for Graph Neural Networks in Blockchain Anomaly Detection
by: Duy, Dang Sy, et al.
Published: (2026) -
Soil Compaction Parameters Prediction Based on Automated Machine Learning Approach
by: Erden, Caner, et al.
Published: (2025) -
An Improved Adaptive PID Optimizer with Enhanced Convergence and Stability for Deep Learning
by: Saini, Saurabh, et al.
Published: (2026)