Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tanada, Kosei, Iwanaga, Yuka, Tsuchinaga, Masayoshi, Nakamura, Yuji, Mori, Takemitsu, Sakai, Remi, Yamamoto, Takashi
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912179133874176
author Tanada, Kosei
Iwanaga, Yuka
Tsuchinaga, Masayoshi
Nakamura, Yuji
Mori, Takemitsu
Sakai, Remi
Yamamoto, Takashi
author_facet Tanada, Kosei
Iwanaga, Yuka
Tsuchinaga, Masayoshi
Nakamura, Yuji
Mori, Takemitsu
Sakai, Remi
Yamamoto, Takashi
contents To use assistive robots in everyday life, a remote control system with common devices, such as 2D devices, is helpful to control the robots anytime and anywhere as intended. Hand-drawn sketches are one of the intuitive ways to control robots with 2D devices. However, since similar sketches have different intentions from scene to scene, existing work needs additional modalities to set the sketches' semantics. This requires complex operations for users and leads to decreasing usability. In this paper, we propose Sketch-MoMa, a teleoperation system using the user-given hand-drawn sketches as instructions to control a robot. We use Vision-Language Models (VLMs) to understand the user-given sketches superimposed on an observation image and infer drawn shapes and low-level tasks of the robot. We utilize the sketches and the generated shapes for recognition and motion planning of the generated low-level tasks for precise and intuitive operations. We validate our approach using state-of-the-art VLMs with 7 tasks and 5 sketch shapes. We also demonstrate that our approach effectively specifies the detailed motions, such as how to grasp and how much to rotate. Moreover, we show the competitive usability of our approach compared with the existing 2D interface through a user experiment with 14 participants.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19153
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches
Tanada, Kosei
Iwanaga, Yuka
Tsuchinaga, Masayoshi
Nakamura, Yuji
Mori, Takemitsu
Sakai, Remi
Yamamoto, Takashi
Robotics
To use assistive robots in everyday life, a remote control system with common devices, such as 2D devices, is helpful to control the robots anytime and anywhere as intended. Hand-drawn sketches are one of the intuitive ways to control robots with 2D devices. However, since similar sketches have different intentions from scene to scene, existing work needs additional modalities to set the sketches' semantics. This requires complex operations for users and leads to decreasing usability. In this paper, we propose Sketch-MoMa, a teleoperation system using the user-given hand-drawn sketches as instructions to control a robot. We use Vision-Language Models (VLMs) to understand the user-given sketches superimposed on an observation image and infer drawn shapes and low-level tasks of the robot. We utilize the sketches and the generated shapes for recognition and motion planning of the generated low-level tasks for precise and intuitive operations. We validate our approach using state-of-the-art VLMs with 7 tasks and 5 sketch shapes. We also demonstrate that our approach effectively specifies the detailed motions, such as how to grasp and how much to rotate. Moreover, we show the competitive usability of our approach compared with the existing 2D interface through a user experiment with 14 participants.
title Sketch-MoMa: Teleoperation for Mobile Manipulator via Interpretation of Hand-Drawn Sketches
topic Robotics
url https://arxiv.org/abs/2412.19153