VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Hung, Kuo-Han, Lo, Pang-Chi, Yeh, Jia-Fong, Hsu, Han-Yuan, Chen, Yi-Ting, Hsu, Winston H.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912238731788288
author Hung, Kuo-Han
Lo, Pang-Chi
Yeh, Jia-Fong
Hsu, Han-Yuan
Chen, Yi-Ting
Hsu, Winston H.
author_facet Hung, Kuo-Han
Lo, Pang-Chi
Yeh, Jia-Fong
Hsu, Han-Yuan
Chen, Yi-Ting
Hsu, Winston H.
contents We study reward models for long-horizon manipulation tasks by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Recent advancements in cross-modality modeling have highlighted the potential of reward modeling through visual and language correlations. However, existing VIC methods face challenges in learning rewards for long-horizon tasks due to their lack of sub-stage awareness, difficulty in modeling task complexities, and inadequate object state estimation. To address these challenges, we introduce VICtoR, a novel hierarchical VIC reward model capable of providing effective reward signals for long-horizon manipulation tasks. VICtoR precisely assesses task progress at various levels through a novel stage detector and motion progress evaluator, offering insightful guidance for agents learning the task effectively. To validate the effectiveness of VICtoR, we conducted extensive experiments in both simulated and real-world environments. The results suggest that VICtoR outperformed the best existing VIC methods, achieving a 43% improvement in success rates for long-horizon tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2405_16545
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation
Hung, Kuo-Han
Lo, Pang-Chi
Yeh, Jia-Fong
Hsu, Han-Yuan
Chen, Yi-Ting
Hsu, Winston H.
Robotics
We study reward models for long-horizon manipulation tasks by learning from action-free videos and language instructions, which we term the visual-instruction correlation (VIC) problem. Recent advancements in cross-modality modeling have highlighted the potential of reward modeling through visual and language correlations. However, existing VIC methods face challenges in learning rewards for long-horizon tasks due to their lack of sub-stage awareness, difficulty in modeling task complexities, and inadequate object state estimation. To address these challenges, we introduce VICtoR, a novel hierarchical VIC reward model capable of providing effective reward signals for long-horizon manipulation tasks. VICtoR precisely assesses task progress at various levels through a novel stage detector and motion progress evaluator, offering insightful guidance for agents learning the task effectively. To validate the effectiveness of VICtoR, we conducted extensive experiments in both simulated and real-world environments. The results suggest that VICtoR outperformed the best existing VIC methods, achieving a 43% improvement in success rates for long-horizon tasks.
title VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation
topic Robotics
url https://arxiv.org/abs/2405.16545