Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kambara, Motonari, Sugiura, Komei
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916555687723008
author Kambara, Motonari
Sugiura, Komei
author_facet Kambara, Motonari
Sugiura, Komei
contents This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based on natural language instructions, egocentric view images before manipulation, and the given end-effector trajectories. Conventional methods typically perform success prediction only after the manipulation is executed, limiting their efficiency in executing the entire task sequence. We propose a novel approach that enables the prediction of success or failure by aligning the given trajectories and images with natural language instructions. We introduce Trajectory Encoder to apply learnable weighting to the input trajectories, allowing the model to consider temporal dynamics and interactions between objects and the end effector, improving the model's ability to predict manipulation outcomes accurately. We constructed a dataset based on the RT-1 dataset, a large-scale benchmark for open-vocabulary object manipulation tasks, to evaluate our method. The experimental results show that our method achieved a higher prediction accuracy than baseline approaches.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19112
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
Kambara, Motonari
Sugiura, Komei
Robotics
Computer Vision and Pattern Recognition
This study addresses a task designed to predict the future success or failure of open-vocabulary object manipulation. In this task, the model is required to make predictions based on natural language instructions, egocentric view images before manipulation, and the given end-effector trajectories. Conventional methods typically perform success prediction only after the manipulation is executed, limiting their efficiency in executing the entire task sequence. We propose a novel approach that enables the prediction of success or failure by aligning the given trajectories and images with natural language instructions. We introduce Trajectory Encoder to apply learnable weighting to the input trajectories, allowing the model to consider temporal dynamics and interactions between objects and the end effector, improving the model's ability to predict manipulation outcomes accurately. We constructed a dataset based on the RT-1 dataset, a large-scale benchmark for open-vocabulary object manipulation tasks, to evaluate our method. The experimental results show that our method achieved a higher prediction accuracy than baseline approaches.
title Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19112