ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Jiahui, Luo, Yusen, Anwar, Abrar, Sontakke, Sumedh Anand, Lim, Joseph J, Thomason, Jesse, Biyik, Erdem, Zhang, Jesse
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915504318316544
author Zhang, Jiahui
Luo, Yusen
Anwar, Abrar
Sontakke, Sumedh Anand
Lim, Joseph J
Thomason, Jesse
Biyik, Erdem
Zhang, Jesse
author_facet Zhang, Jiahui
Luo, Yusen
Anwar, Abrar
Sontakke, Sumedh Anand
Lim, Joseph J
Thomason, Jesse
Biyik, Erdem
Zhang, Jesse
contents We introduce ReWiND, a framework for learning robot manipulation tasks solely from language instructions without per-task demonstrations. Standard reinforcement learning (RL) and imitation learning methods require expert supervision through human-designed reward functions or demonstrations for every new task. In contrast, ReWiND starts from a small demonstration dataset to learn: (1) a data-efficient, language-conditioned reward function that labels the dataset with rewards, and (2) a language-conditioned policy pre-trained with offline RL using these rewards. Given an unseen task variation, ReWiND fine-tunes the pre-trained policy using the learned reward function, requiring minimal online interaction. We show that ReWiND's reward model generalizes effectively to unseen tasks, outperforming baselines by up to 2.4x in reward generalization and policy alignment metrics. Finally, we demonstrate that ReWiND enables sample-efficient adaptation to new tasks, beating baselines by 2x in simulation and improving real-world pretrained bimanual policies by 5x, taking a step towards scalable, real-world robot learning. See website at https://rewind-reward.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2505_10911
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
Zhang, Jiahui
Luo, Yusen
Anwar, Abrar
Sontakke, Sumedh Anand
Lim, Joseph J
Thomason, Jesse
Biyik, Erdem
Zhang, Jesse
Robotics
We introduce ReWiND, a framework for learning robot manipulation tasks solely from language instructions without per-task demonstrations. Standard reinforcement learning (RL) and imitation learning methods require expert supervision through human-designed reward functions or demonstrations for every new task. In contrast, ReWiND starts from a small demonstration dataset to learn: (1) a data-efficient, language-conditioned reward function that labels the dataset with rewards, and (2) a language-conditioned policy pre-trained with offline RL using these rewards. Given an unseen task variation, ReWiND fine-tunes the pre-trained policy using the learned reward function, requiring minimal online interaction. We show that ReWiND's reward model generalizes effectively to unseen tasks, outperforming baselines by up to 2.4x in reward generalization and policy alignment metrics. Finally, we demonstrate that ReWiND enables sample-efficient adaptation to new tasks, beating baselines by 2x in simulation and improving real-world pretrained bimanual policies by 5x, taking a step towards scalable, real-world robot learning. See website at https://rewind-reward.github.io/.
title ReWiND: Language-Guided Rewards Teach Robot Policies without New Demonstrations
topic Robotics
url https://arxiv.org/abs/2505.10911