ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Nan, Pang, Jing-Cheng, Li, Guanlin, Qian, Chao, Yu, Yang
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915516626501632
author Tang, Nan
Pang, Jing-Cheng
Li, Guanlin
Qian, Chao
Yu, Yang
author_facet Tang, Nan
Pang, Jing-Cheng
Li, Guanlin
Qian, Chao
Yu, Yang
contents Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on the distance to a target position. However, such precise positional information is often unavailable in real-world visual settings due to sensory and perceptual limitations. In this study, we propose a method that implicitly infers spatial distances through keypoints extracted from images. Building on this, we introduce Reward Learning with Anticipation Model (ReLAM), a novel framework that automatically generates dense, structured rewards from action-free video demonstrations. ReLAM first learns an anticipation model that serves as a planner and proposes intermediate keypoint-based subgoals on the optimal path to the final goal, creating a structured learning curriculum directly aligned with the task's geometric objectives. Based on the anticipated subgoals, a continuous reward signal is provided to train a low-level, goal-conditioned policy under the hierarchical reinforcement learning (HRL) framework with provable sub-optimality bound. Extensive experiments on complex, long-horizon manipulation tasks show that ReLAM significantly accelerates learning and achieves superior performance compared to state-of-the-art methods.
format Preprint
id arxiv_https___arxiv_org_abs_2509_22402
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
Tang, Nan
Pang, Jing-Cheng
Li, Guanlin
Qian, Chao
Yu, Yang
Machine Learning
Robotics
Reward design remains a critical bottleneck in visual reinforcement learning (RL) for robotic manipulation. In simulated environments, rewards are conventionally designed based on the distance to a target position. However, such precise positional information is often unavailable in real-world visual settings due to sensory and perceptual limitations. In this study, we propose a method that implicitly infers spatial distances through keypoints extracted from images. Building on this, we introduce Reward Learning with Anticipation Model (ReLAM), a novel framework that automatically generates dense, structured rewards from action-free video demonstrations. ReLAM first learns an anticipation model that serves as a planner and proposes intermediate keypoint-based subgoals on the optimal path to the final goal, creating a structured learning curriculum directly aligned with the task's geometric objectives. Based on the anticipated subgoals, a continuous reward signal is provided to train a low-level, goal-conditioned policy under the hierarchical reinforcement learning (HRL) framework with provable sub-optimality bound. Extensive experiments on complex, long-horizon manipulation tasks show that ReLAM significantly accelerates learning and achieves superior performance compared to state-of-the-art methods.
title ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
topic Machine Learning
Robotics
url https://arxiv.org/abs/2509.22402