Semi-gradient DICE for Offline Constrained Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Woosung, Seo, JunHo, Lee, Jongmin, Lee, Byung-Jun
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913887911149568
author Kim, Woosung
Seo, JunHo
Lee, Jongmin
Lee, Byung-Jun
author_facet Kim, Woosung
Seo, JunHo
Lee, Jongmin
Lee, Byung-Jun
contents Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliable off-policy evaluation (OPE) and policy optimization. DICE-based offline constrained RL particularly benefits from the flexibility of DICE, as it simultaneously maximizes return while estimating costs in offline settings. However, we have observed that recent approaches designed to enhance the offline RL performance of the DICE framework inadvertently undermine its ability to perform OPE, making them unsuitable for constrained RL scenarios. In this paper, we identify the root cause of this limitation: their reliance on a semi-gradient optimization, which solves a fundamentally different optimization problem and results in failures in cost estimation. Building on these insights, we propose a novel method to enable OPE and constrained RL through semi-gradient DICE. Our method ensures accurate cost estimation and achieves state-of-the-art performance on the offline constrained RL benchmark, DSRL.
format Preprint
id arxiv_https___arxiv_org_abs_2506_08644
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Semi-gradient DICE for Offline Constrained Reinforcement Learning
Kim, Woosung
Seo, JunHo
Lee, Jongmin
Lee, Byung-Jun
Machine Learning
Stationary Distribution Correction Estimation (DICE) addresses the mismatch between the stationary distribution induced by a policy and the target distribution required for reliable off-policy evaluation (OPE) and policy optimization. DICE-based offline constrained RL particularly benefits from the flexibility of DICE, as it simultaneously maximizes return while estimating costs in offline settings. However, we have observed that recent approaches designed to enhance the offline RL performance of the DICE framework inadvertently undermine its ability to perform OPE, making them unsuitable for constrained RL scenarios. In this paper, we identify the root cause of this limitation: their reliance on a semi-gradient optimization, which solves a fundamentally different optimization problem and results in failures in cost estimation. Building on these insights, we propose a novel method to enable OPE and constrained RL through semi-gradient DICE. Our method ensures accurate cost estimation and achieves state-of-the-art performance on the offline constrained RL benchmark, DSRL.
title Semi-gradient DICE for Offline Constrained Reinforcement Learning
topic Machine Learning
url https://arxiv.org/abs/2506.08644