Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bhatt, Dwait, Chou, Shih-Chieh, Atanasov, Nikolay
Natura: Preprint
Pubblicazione: 2026
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908919096410112
author Bhatt, Dwait
Chou, Shih-Chieh
Atanasov, Nikolay
author_facet Bhatt, Dwait
Chou, Shih-Chieh
Atanasov, Nikolay
contents Several approaches have been proposed to improve the sample efficiency of online reinforcement learning (RL) by leveraging demonstrations collected offline. The offline data can be used directly as transitions to optimize RL objectives, or offline policy and value functions can first be learned from the data and then used for online finetuning or to provide reference actions. While each of these strategies has shown compelling results, it is unclear which method has the most impact on sample efficiency, whether these approaches can be combined, and if there are cumulative benefits. We classify existing demonstration-augmented RL approaches into three categories and perform an extensive empirical study of their strengths, weaknesses, and combinations to isolate the contribution of each strategy and determine effective hybrid combinations for sample-efficient online RL. Our analysis reveals that directly reusing offline data and initializing with behavior cloning consistently outperform more complex offline RL pretraining methods for improving online sample efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2603_27400
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning
Bhatt, Dwait
Chou, Shih-Chieh
Atanasov, Nikolay
Robotics
Machine Learning
Several approaches have been proposed to improve the sample efficiency of online reinforcement learning (RL) by leveraging demonstrations collected offline. The offline data can be used directly as transitions to optimize RL objectives, or offline policy and value functions can first be learned from the data and then used for online finetuning or to provide reference actions. While each of these strategies has shown compelling results, it is unclear which method has the most impact on sample efficiency, whether these approaches can be combined, and if there are cumulative benefits. We classify existing demonstration-augmented RL approaches into three categories and perform an extensive empirical study of their strengths, weaknesses, and combinations to isolate the contribution of each strategy and determine effective hybrid combinations for sample-efficient online RL. Our analysis reveals that directly reusing offline data and initializing with behavior cloning consistently outperform more complex offline RL pretraining methods for improving online sample efficiency.
title Rainbow-DemoRL: Combining Improvements in Demonstration-Augmented Reinforcement Learning
topic Robotics
Machine Learning
url https://arxiv.org/abs/2603.27400