Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Kevin, Scalise, Rosario, Winston, Cleah, Agrawal, Ayush, Zhang, Yunchu, Baijal, Rohan, Grotz, Markus, Boots, Byron, Burchfiel, Benjamin, Itkina, Masha, Shah, Paarth, Gupta, Abhishek
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909869464879104
author Huang, Kevin
Scalise, Rosario
Winston, Cleah
Agrawal, Ayush
Zhang, Yunchu
Baijal, Rohan
Grotz, Markus
Boots, Byron
Burchfiel, Benjamin
Itkina, Masha
Shah, Paarth
Gupta, Abhishek
author_facet Huang, Kevin
Scalise, Rosario
Winston, Cleah
Agrawal, Ayush
Zhang, Yunchu
Baijal, Rohan
Grotz, Markus
Boots, Byron
Burchfiel, Benjamin
Itkina, Masha
Shah, Paarth
Gupta, Abhishek
contents Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In contrast, non-expert data -- such as play data, suboptimal demonstrations, partial task completions, or rollouts from suboptimal policies -- can offer broader coverage and lower collection costs. However, conventional imitation learning approaches fail to utilize this data effectively. To address these challenges, we posit that with right design decisions, offline reinforcement learning can be used as a tool to harness non-expert data to enhance the performance of imitation learning policies. We show that while standard offline RL approaches can be ineffective at actually leveraging non-expert data under the sparse data coverage settings typically encountered in the real world, simple algorithmic modifications can allow for the utilization of this data, without significant additional assumptions. Our approach shows that broadening the support of the policy distribution can allow imitation algorithms augmented by offline RL to solve tasks robustly, showing considerably enhanced recovery and generalization behavior. In manipulation tasks, these innovations significantly increase the range of initial conditions where learned policies are successful when non-expert data is incorporated. Moreover, we show that these methods are able to leverage all collected data, including partial or suboptimal demonstrations, to bolster task-directed policy performance. This underscores the importance of algorithmic techniques for using non-expert data for robust policy learning in robotics. Website: https://uwrobotlearning.github.io/RISE-offline/
format Preprint
id arxiv_https___arxiv_org_abs_2510_19495
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
Huang, Kevin
Scalise, Rosario
Winston, Cleah
Agrawal, Ayush
Zhang, Yunchu
Baijal, Rohan
Grotz, Markus
Boots, Byron
Burchfiel, Benjamin
Itkina, Masha
Shah, Paarth
Gupta, Abhishek
Robotics
Artificial Intelligence
Machine Learning
Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In contrast, non-expert data -- such as play data, suboptimal demonstrations, partial task completions, or rollouts from suboptimal policies -- can offer broader coverage and lower collection costs. However, conventional imitation learning approaches fail to utilize this data effectively. To address these challenges, we posit that with right design decisions, offline reinforcement learning can be used as a tool to harness non-expert data to enhance the performance of imitation learning policies. We show that while standard offline RL approaches can be ineffective at actually leveraging non-expert data under the sparse data coverage settings typically encountered in the real world, simple algorithmic modifications can allow for the utilization of this data, without significant additional assumptions. Our approach shows that broadening the support of the policy distribution can allow imitation algorithms augmented by offline RL to solve tasks robustly, showing considerably enhanced recovery and generalization behavior. In manipulation tasks, these innovations significantly increase the range of initial conditions where learned policies are successful when non-expert data is incorporated. Moreover, we show that these methods are able to leverage all collected data, including partial or suboptimal demonstrations, to bolster task-directed policy performance. This underscores the importance of algorithmic techniques for using non-expert data for robust policy learning in robotics. Website: https://uwrobotlearning.github.io/RISE-offline/
title Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.19495