COOPO: Cyclic Offline-Online Policy Optimization Algorithm

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Qisai, Jiang, Zhanhong, Waite, Joshua Russell, Balu, Aditya, Fleming, Cody, Sarkar, Soumik
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910232918097920
author Liu, Qisai
Jiang, Zhanhong
Waite, Joshua Russell
Balu, Aditya
Fleming, Cody
Sarkar, Soumik
author_facet Liu, Qisai
Jiang, Zhanhong
Waite, Joshua Russell
Balu, Aditya
Fleming, Cody
Sarkar, Soumik
contents Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment interactions. The recent advent of hybrid offline-to-online methods bridges these domains but suffers from distribution drift during transitions and catastrophic forgetting of offline knowledge. We introduce COOPO (Cyclic Offline-Online Policy Optimization), a generalized framework that repeatedly cycles between constrained offline training and online fine-tuning. Each cycle first anchors the policy to the dataset via KL-regularized advantage-weighted offline updates to minimize distributional shift and then fine-tunes it online using any policy optimization for stable exploration. Crucially, periodically returning to offline training eliminates forgetting and drift while maximizing dataset reuse. The cyclic behavior also helps reduce the online environment interactions. Theoretically, COOPO achieves better online sample efficiency, surpassing pure online RL, with guaranteed monotonic improvement under standard coverage assumptions. Extensive D4RL benchmarks demonstrate COOPO reduces online interactions versus state-of-the-art hybrids while improving final returns, maintaining robustness across diverse offline algorithms and online optimizers. This looped synergy sets new efficiency and performance standards for adaptive RL.
format Preprint
id arxiv_https___arxiv_org_abs_2605_18675
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle COOPO: Cyclic Offline-Online Policy Optimization Algorithm
Liu, Qisai
Jiang, Zhanhong
Waite, Joshua Russell
Balu, Aditya
Fleming, Cody
Sarkar, Soumik
Machine Learning
Artificial Intelligence
Offline reinforcement learning struggles with distributional shift and constrained performance due to static dataset limitations, while online RL demands prohibitive environment interactions. The recent advent of hybrid offline-to-online methods bridges these domains but suffers from distribution drift during transitions and catastrophic forgetting of offline knowledge. We introduce COOPO (Cyclic Offline-Online Policy Optimization), a generalized framework that repeatedly cycles between constrained offline training and online fine-tuning. Each cycle first anchors the policy to the dataset via KL-regularized advantage-weighted offline updates to minimize distributional shift and then fine-tunes it online using any policy optimization for stable exploration. Crucially, periodically returning to offline training eliminates forgetting and drift while maximizing dataset reuse. The cyclic behavior also helps reduce the online environment interactions. Theoretically, COOPO achieves better online sample efficiency, surpassing pure online RL, with guaranteed monotonic improvement under standard coverage assumptions. Extensive D4RL benchmarks demonstrate COOPO reduces online interactions versus state-of-the-art hybrids while improving final returns, maintaining robustness across diverse offline algorithms and online optimizers. This looped synergy sets new efficiency and performance standards for adaptive RL.
title COOPO: Cyclic Offline-Online Policy Optimization Algorithm
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2605.18675