Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kong, Lingkai, Wang, Haichuan, Wang, Tonghan, Xiong, Guojun, Tambe, Milind
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908779865440256
author Kong, Lingkai
Wang, Haichuan
Wang, Tonghan
Xiong, Guojun
Tambe, Milind
author_facet Kong, Lingkai
Wang, Haichuan
Wang, Tonghan
Xiong, Guojun
Tambe, Milind
contents Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamics in the offline dataset differ from those encountered online. Existing approaches typically mitigate this issue by penalizing or filtering offline transitions in regions with large dynamics gap. However, their dynamics-gap estimators often rely on KL divergence or mutual information, which can be ill-defined when offline and online dynamics have mismatched support. To address this challenge, we propose CompFlow, a principled framework built on the theoretical connection between flow matching and optimal transport. Specifically, we model the online dynamics as a conditional flow built upon the output distribution of a pretrained offline flow, rather than learning it directly from a Gaussian prior. This composite structure provides two advantages: (1) improved generalization when learning online dynamics under limited interaction data, and (2) a well-defined and stable estimate of the dynamics gap via the Wasserstein distance between offline and online transitions. Building on this dynamics-gap estimator, we further develop an optimistic active data collection strategy that prioritizes exploration in high-gap regions, and show theoretically that it reduces the performance gap to the optimal policy. Empirically, CompFlow consistently outperforms strong baselines across a range of RL benchmarks with shifted-dynamics data.
format Preprint
id arxiv_https___arxiv_org_abs_2505_23062
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
Kong, Lingkai
Wang, Haichuan
Wang, Tonghan
Xiong, Guojun
Tambe, Milind
Machine Learning
Artificial Intelligence
Incorporating pre-collected offline data can substantially improve the sample efficiency of reinforcement learning (RL), but its benefits can break down when the transition dynamics in the offline dataset differ from those encountered online. Existing approaches typically mitigate this issue by penalizing or filtering offline transitions in regions with large dynamics gap. However, their dynamics-gap estimators often rely on KL divergence or mutual information, which can be ill-defined when offline and online dynamics have mismatched support. To address this challenge, we propose CompFlow, a principled framework built on the theoretical connection between flow matching and optimal transport. Specifically, we model the online dynamics as a conditional flow built upon the output distribution of a pretrained offline flow, rather than learning it directly from a Gaussian prior. This composite structure provides two advantages: (1) improved generalization when learning online dynamics under limited interaction data, and (2) a well-defined and stable estimate of the dynamics gap via the Wasserstein distance between offline and online transitions. Building on this dynamics-gap estimator, we further develop an optimistic active data collection strategy that prioritizes exploration in high-gap regions, and show theoretically that it reduces the performance gap to the optimal policy. Empirically, CompFlow consistently outperforms strong baselines across a range of RL benchmarks with shifted-dynamics data.
title Composite Flow Matching for Reinforcement Learning with Shifted-Dynamics Data
topic Machine Learning
Artificial Intelligence
url https://arxiv.org/abs/2505.23062