Dynamical Priors as a Training Objective in Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Author: Subaharan, Sukesh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918464244940800
author Subaharan, Sukesh
author_facet Subaharan, Sukesh
contents Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policies may achieve high performance while exhibiting temporally incoherent behavior such as abrupt confidence shifts, oscillations, or degenerate inactivity. We introduce Dynamical Prior Reinforcement Learning (DP-RL), a training framework that augments policy gradient learning with an auxiliary loss derived from external state dynamics that implement evidence accumulation and hysteresis. Without modifying the reward, environment, or policy architecture, this prior shapes the temporal evolution of action probabilities during learning. Across three minimal environments, we show that dynamical priors systematically alter decision trajectories in task-dependent ways, promoting temporally structured behavior that cannot be explained by generic smoothing. These results demonstrate that training objectives alone can control the temporal geometry of decision-making in RL agents.
format Preprint
id arxiv_https___arxiv_org_abs_2604_21464
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Dynamical Priors as a Training Objective in Reinforcement Learning
Subaharan, Sukesh
Machine Learning
Artificial Intelligence
68T05
I.2.6; I.2.11
Standard reinforcement learning (RL) optimizes policies for reward but imposes few constraints on how decisions evolve over time. As a result, policies may achieve high performance while exhibiting temporally incoherent behavior such as abrupt confidence shifts, oscillations, or degenerate inactivity. We introduce Dynamical Prior Reinforcement Learning (DP-RL), a training framework that augments policy gradient learning with an auxiliary loss derived from external state dynamics that implement evidence accumulation and hysteresis. Without modifying the reward, environment, or policy architecture, this prior shapes the temporal evolution of action probabilities during learning. Across three minimal environments, we show that dynamical priors systematically alter decision trajectories in task-dependent ways, promoting temporally structured behavior that cannot be explained by generic smoothing. These results demonstrate that training objectives alone can control the temporal geometry of decision-making in RL agents.
title Dynamical Priors as a Training Objective in Reinforcement Learning
topic Machine Learning
Artificial Intelligence
68T05
I.2.6; I.2.11
url https://arxiv.org/abs/2604.21464