Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mahesheka, Harsh, Xie, Zhixian, Wang, Zhaoran, Jin, Wanxin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914970462060544
author Mahesheka, Harsh
Xie, Zhixian
Wang, Zhaoran
Jin, Wanxin
author_facet Mahesheka, Harsh
Xie, Zhixian
Wang, Zhaoran
Jin, Wanxin
contents Learning from Demonstrations, particularly from biological experts like humans and animals, often encounters significant data acquisition challenges. While recent approaches leverage internet videos for learning, they require complex, task-specific pipelines to extract and retarget motion data for the agent. In this work, we introduce a language-model-assisted bi-level programming framework that enables a reinforcement learning agent to directly learn its reward from internet videos, bypassing dedicated data preparation. The framework includes two levels: an upper level where a vision-language model (VLM) provides feedback by comparing the learner's behavior with expert videos, and a lower level where a large language model (LLM) translates this feedback into reward updates. The VLM and LLM collaborate within this bi-level framework, using a "chain rule" approach to derive a valid search direction for reward learning. We validate the method for reward learning from YouTube videos, and the results have shown that the proposed method enables efficient reward design from expert videos of biological agents for complex behavior synthesis.
format Preprint
id arxiv_https___arxiv_org_abs_2410_09286
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
Mahesheka, Harsh
Xie, Zhixian
Wang, Zhaoran
Jin, Wanxin
Robotics
Artificial Intelligence
Learning from Demonstrations, particularly from biological experts like humans and animals, often encounters significant data acquisition challenges. While recent approaches leverage internet videos for learning, they require complex, task-specific pipelines to extract and retarget motion data for the agent. In this work, we introduce a language-model-assisted bi-level programming framework that enables a reinforcement learning agent to directly learn its reward from internet videos, bypassing dedicated data preparation. The framework includes two levels: an upper level where a vision-language model (VLM) provides feedback by comparing the learner's behavior with expert videos, and a lower level where a large language model (LLM) translates this feedback into reward updates. The VLM and LLM collaborate within this bi-level framework, using a "chain rule" approach to derive a valid search direction for reward learning. We validate the method for reward learning from YouTube videos, and the results have shown that the proposed method enables efficient reward design from expert videos of biological agents for complex behavior synthesis.
title Language-Model-Assisted Bi-Level Programming for Reward Learning from Internet Videos
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2410.09286