Distributions as Actions: A Unified Framework for Diverse Action Spaces

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: He, Jiamin, Mahmood, A. Rupam, White, Martha
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917492391149568
author He, Jiamin
Mahmood, A. Rupam
White, Martha
author_facet He, Jiamin
Mahmood, A. Rupam
White, Martha
contents We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary between agent and environment. This reparameterization makes the new action space continuous, regardless of the original action type (discrete, continuous, hybrid, etc.). Under this new parameterization, we develop a generalized deterministic policy gradient estimator, Distributions-as-Actions Policy Gradient (DA-PG), which has lower variance than the gradient in the original action space. Although learning the critic over distribution parameters poses new challenges, we introduce Interpolated Critic Learning (ICL), a simple yet effective strategy to enhance learning, supported by insights from bandit settings. Building on TD3, a strong baseline for continuous control, we propose a practical actor-critic algorithm, Distributions-as-Actions Actor-Critic (DA-AC). Empirically, DA-AC achieves competitive performance in various settings across discrete, continuous, and hybrid control.
format Preprint
id arxiv_https___arxiv_org_abs_2506_16608
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Distributions as Actions: A Unified Framework for Diverse Action Spaces
He, Jiamin
Mahmood, A. Rupam
White, Martha
Machine Learning
Artificial Intelligence
We introduce a novel reinforcement learning (RL) framework that treats parameterized action distributions as actions, redefining the boundary between agent and environment. This reparameterization makes the new action space continuous, regardless of the original action type (discrete, continuous, hybrid, etc.). Under this new parameterization, we develop a generalized deterministic policy gradient estimator, Distributions-as-Actions Policy Gradient (DA-PG), which has lower variance than the gradient in the original action space. Although learning the critic over distribution parameters poses new challenges, we introduce Interpolated Critic Learning (ICL), a simple yet effective strategy to enhance learning, supported by insights from bandit settings. Building on TD3, a strong baseline for continuous control, we propose a practical actor-critic algorithm, Distributions-as-Actions Actor-Critic (DA-AC). Empirically, DA-AC achieves competitive performance in various settings across discrete, continuous, and hybrid control.
title Distributions as Actions: A Unified Framework for Diverse Action Spaces
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2506.16608