GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wake, Naoki, Kanehira, Atsushi, Sasabuchi, Kazuhiro, Takamatsu, Jun, Ikeuchi, Katsushi
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912065480818688
author Wake, Naoki
Kanehira, Atsushi
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
author_facet Wake, Naoki
Kanehira, Atsushi
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
contents We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable robot programs that incorporate insights into affordances. The process begins with GPT-4V analyzing the videos to obtain textual explanations of environmental and action details. A GPT-4-based task planner then encodes these details into a symbolic task plan. Subsequently, vision systems spatially and temporally ground the task plan in the videos. Objects are identified using an open-vocabulary object detector, and hand-object interactions are analyzed to pinpoint moments of grasping and releasing. This spatiotemporal grounding allows for the gathering of affordance information (e.g., grasp types, waypoints, and body postures) critical for robot execution. Experiments across various scenarios demonstrate the method's efficacy in enabling real robots to operate from one-shot human demonstrations. Meanwhile, quantitative tests have revealed instances of hallucination in GPT-4V, highlighting the importance of incorporating human supervision within the pipeline. The prompts of GPT-4V/GPT-4 are available at this project page: https://microsoft.github.io/GPT4Vision-Robot-Manipulation-Prompts/
format Preprint
id arxiv_https___arxiv_org_abs_2311_12015
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
Wake, Naoki
Kanehira, Atsushi
Sasabuchi, Kazuhiro
Takamatsu, Jun
Ikeuchi, Katsushi
Robotics
Computation and Language
Computer Vision and Pattern Recognition
We introduce a pipeline that enhances a general-purpose Vision Language Model, GPT-4V(ision), to facilitate one-shot visual teaching for robotic manipulation. This system analyzes videos of humans performing tasks and outputs executable robot programs that incorporate insights into affordances. The process begins with GPT-4V analyzing the videos to obtain textual explanations of environmental and action details. A GPT-4-based task planner then encodes these details into a symbolic task plan. Subsequently, vision systems spatially and temporally ground the task plan in the videos. Objects are identified using an open-vocabulary object detector, and hand-object interactions are analyzed to pinpoint moments of grasping and releasing. This spatiotemporal grounding allows for the gathering of affordance information (e.g., grasp types, waypoints, and body postures) critical for robot execution. Experiments across various scenarios demonstrate the method's efficacy in enabling real robots to operate from one-shot human demonstrations. Meanwhile, quantitative tests have revealed instances of hallucination in GPT-4V, highlighting the importance of incorporating human supervision within the pipeline. The prompts of GPT-4V/GPT-4 are available at this project page: https://microsoft.github.io/GPT4Vision-Robot-Manipulation-Prompts/
title GPT-4V(ision) for Robotics: Multimodal Task Planning from Human Demonstration
topic Robotics
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.12015