Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qiao, Zhongjian, Yang, Rui, Lyu, Jiafei, Li, Xiu, Dai, Zhongxiang, Yang, Zhuoran, Gao, Siyang, Qiu, Shuang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912953451675648
author Qiao, Zhongjian
Yang, Rui
Lyu, Jiafei
Li, Xiu
Dai, Zhongxiang
Yang, Zhuoran
Gao, Siyang
Qiu, Shuang
author_facet Qiao, Zhongjian
Yang, Rui
Lyu, Jiafei
Li, Xiu
Dai, Zhongxiang
Yang, Zhuoran
Gao, Siyang
Qiu, Shuang
contents Single-domain offline reinforcement learning (RL) often suffers from limited data coverage, while cross-domain offline RL handles this issue by leveraging additional data from other domains with dynamics shifts. However, existing studies primarily focus on train-time robustness (handling dynamics shifts from training data), neglecting the test-time robustness against dynamics perturbations when deployed in practical scenarios. In this paper, we investigate dual (both train-time and test-time) robustness against dynamics shifts in cross-domain offline RL. We first empirically show that the policy trained with cross-domain offline RL exhibits fragility under dynamics perturbations during evaluation, particularly when target domain data is limited. To address this, we introduce a novel robust cross-domain Bellman (RCB) operator, which enhances test-time robustness against dynamics perturbations while staying conservative to the out-of-distribution dynamics transitions, thus guaranteeing the train-time robustness. To further counteract potential value overestimation or underestimation caused by the RCB operator, we introduce two techniques, the dynamic value penalty and the Huber loss, into our framework, resulting in the practical \textbf{D}ual-\textbf{RO}bust \textbf{C}ross-domain \textbf{O}ffline RL (DROCO) algorithm. Extensive empirical results across various dynamics shift scenarios show that DROCO outperforms strong baselines and exhibits enhanced robustness to dynamics perturbations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02486
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
Qiao, Zhongjian
Yang, Rui
Lyu, Jiafei
Li, Xiu
Dai, Zhongxiang
Yang, Zhuoran
Gao, Siyang
Qiu, Shuang
Machine Learning
Single-domain offline reinforcement learning (RL) often suffers from limited data coverage, while cross-domain offline RL handles this issue by leveraging additional data from other domains with dynamics shifts. However, existing studies primarily focus on train-time robustness (handling dynamics shifts from training data), neglecting the test-time robustness against dynamics perturbations when deployed in practical scenarios. In this paper, we investigate dual (both train-time and test-time) robustness against dynamics shifts in cross-domain offline RL. We first empirically show that the policy trained with cross-domain offline RL exhibits fragility under dynamics perturbations during evaluation, particularly when target domain data is limited. To address this, we introduce a novel robust cross-domain Bellman (RCB) operator, which enhances test-time robustness against dynamics perturbations while staying conservative to the out-of-distribution dynamics transitions, thus guaranteeing the train-time robustness. To further counteract potential value overestimation or underestimation caused by the RCB operator, we introduce two techniques, the dynamic value penalty and the Huber loss, into our framework, resulting in the practical \textbf{D}ual-\textbf{RO}bust \textbf{C}ross-domain \textbf{O}ffline RL (DROCO) algorithm. Extensive empirical results across various dynamics shift scenarios show that DROCO outperforms strong baselines and exhibits enhanced robustness to dynamics perturbations.
title Dual-Robust Cross-Domain Offline Reinforcement Learning Against Dynamics Shifts
topic Machine Learning
url https://arxiv.org/abs/2512.02486