Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Abolfazli, Amir, Song, Zekun, Anand, Avishek, Nejdl, Wolfgang
Format:	Preprint
Published:	2025
Subjects:	Machine Learning
Online Access:	https://arxiv.org/abs/2502.00601
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866909577843310592
author	Abolfazli, Amir Song, Zekun Anand, Avishek Nejdl, Wolfgang
author_facet	Abolfazli, Amir Song, Zekun Anand, Avishek Nejdl, Wolfgang
contents	The success of deep reinforcement learning (DRL) relies on the availability and quality of training data, often requiring extensive interactions with specific environments. In many real-world scenarios, where data collection is costly and risky, offline reinforcement learning (RL) offers a solution by utilizing data collected by domain experts and searching for a batch-constrained optimal policy. This approach is further augmented by incorporating external data sources, expanding the range and diversity of data collection possibilities. However, existing offline RL methods often struggle with challenges posed by non-matching data from these external sources. In this work, we specifically address the problem of source-target domain mismatch in scenarios involving mixed datasets, characterized by a predominance of source data generated from random or suboptimal policies and a limited amount of target data generated from higher-quality policies. To tackle this problem, we introduce Transition Scoring (TS), a novel method that assigns scores to transitions based on their similarity to the target domain, and propose Curriculum Learning-Based Trajectory Valuation (CLTV), which effectively leverages these transition scores to identify and prioritize high-quality trajectories through a curriculum learning approach. Our extensive experiments across various offline RL methods and MuJoCo environments, complemented by rigorous theoretical analysis, demonstrate that CLTV enhances the overall performance and transferability of policies learned by offline RL algorithms.
format	Preprint
id	arxiv_https___arxiv_org_abs_2502_00601
institution	arXiv
publishDate	2025
record_format	arxiv
spellingShingle	Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation Abolfazli, Amir Song, Zekun Anand, Avishek Nejdl, Wolfgang Machine Learning The success of deep reinforcement learning (DRL) relies on the availability and quality of training data, often requiring extensive interactions with specific environments. In many real-world scenarios, where data collection is costly and risky, offline reinforcement learning (RL) offers a solution by utilizing data collected by domain experts and searching for a batch-constrained optimal policy. This approach is further augmented by incorporating external data sources, expanding the range and diversity of data collection possibilities. However, existing offline RL methods often struggle with challenges posed by non-matching data from these external sources. In this work, we specifically address the problem of source-target domain mismatch in scenarios involving mixed datasets, characterized by a predominance of source data generated from random or suboptimal policies and a limited amount of target data generated from higher-quality policies. To tackle this problem, we introduce Transition Scoring (TS), a novel method that assigns scores to transitions based on their similarity to the target domain, and propose Curriculum Learning-Based Trajectory Valuation (CLTV), which effectively leverages these transition scores to identify and prioritize high-quality trajectories through a curriculum learning approach. Our extensive experiments across various offline RL methods and MuJoCo environments, complemented by rigorous theoretical analysis, demonstrate that CLTV enhances the overall performance and transferability of policies learned by offline RL algorithms.
title	Enhancing Offline Reinforcement Learning with Curriculum Learning-Based Trajectory Valuation
topic	Machine Learning
url	https://arxiv.org/abs/2502.00601

Similar Items