TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuyang, Wen, Chuan, Hu, Yihang, Jayaraman, Dinesh, Gao, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914580996816896
author Liu, Yuyang
Wen, Chuan
Hu, Yihang
Jayaraman, Dinesh
Gao, Yang
author_facet Liu, Yuyang
Wen, Chuan
Hu, Yihang
Jayaraman, Dinesh
Gao, Yang
contents Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies the degree to which actions advance the system toward task completion over time. We present TimeRewarder, a simple yet effective reward learning method that derives progress estimation signals from passive videos, including robot demonstrations and human videos, by modeling temporal distances between frame pairs. We then demonstrate how TimeRewarder can supply step-wise proxy rewards to guide reinforcement learning. In our comprehensive experiments on ten challenging Meta-World tasks, we show that TimeRewarder dramatically improves RL for sparse-reward tasks, achieving nearly perfect success in 9/10 tasks with only 200,000 environment interactions per task. This approach outperformed previous methods and even the manually designed environment dense reward on both the final success rate and sample efficiency. Moreover, we show that TimeRewarder pretraining can exploit real-world human videos, highlighting its potential as a scalable approach to rich reward signals from diverse video sources.
format Preprint
id arxiv_https___arxiv_org_abs_2509_26627
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
Liu, Yuyang
Wen, Chuan
Hu, Yihang
Jayaraman, Dinesh
Gao, Yang
Artificial Intelligence
Machine Learning
Robotics
Designing dense rewards is crucial for reinforcement learning (RL), yet in robotics it often demands extensive manual effort and lacks scalability. One promising solution is to view task progress as a dense reward signal, as it quantifies the degree to which actions advance the system toward task completion over time. We present TimeRewarder, a simple yet effective reward learning method that derives progress estimation signals from passive videos, including robot demonstrations and human videos, by modeling temporal distances between frame pairs. We then demonstrate how TimeRewarder can supply step-wise proxy rewards to guide reinforcement learning. In our comprehensive experiments on ten challenging Meta-World tasks, we show that TimeRewarder dramatically improves RL for sparse-reward tasks, achieving nearly perfect success in 9/10 tasks with only 200,000 environment interactions per task. This approach outperformed previous methods and even the manually designed environment dense reward on both the final success rate and sample efficiency. Moreover, we show that TimeRewarder pretraining can exploit real-world human videos, highlighting its potential as a scalable approach to rich reward signals from diverse video sources.
title TimeRewarder: Learning Dense Reward from Passive Videos via Frame-wise Temporal Distance
topic Artificial Intelligence
Machine Learning
Robotics
url https://arxiv.org/abs/2509.26627