Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human Handover

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Jiangshan, Dong, Wenlong, Wang, Jiankun, Meng, Max Q. -H.
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914957820428288
author Liu, Jiangshan
Dong, Wenlong
Wang, Jiankun
Meng, Max Q. -H.
author_facet Liu, Jiangshan
Dong, Wenlong
Wang, Jiankun
Meng, Max Q. -H.
contents Human-robot interaction (HRI) encompasses a wide range of collaborative tasks, with handover being one of the most fundamental. As robots become more integrated into human environments, the potential for service robots to assist in handing objects to humans is increasingly promising. In robot-to-human (R2H) handover, selecting the optimal grasp is crucial for success, as it requires avoiding interference with the humans preferred grasp region and minimizing intrusion into their workspace. Existing methods either inadequately consider geometric information or rely on data-driven approaches, which often struggle to generalize across diverse objects. To address these limitations, we propose a novel zero-shot system that combines semantic and geometric information to generate optimal handover grasps. Our method first identifies grasp regions using semantic knowledge from vision-language models (VLMs) and, by incorporating customized visual prompts, achieves finer granularity in region grounding. A grasp is then selected based on grasp distance and approach angle to maximize human ease and avoid interference. We validate our approach through ablation studies and real-world comparison experiments. Results demonstrate that our system improves handover success rates and provides a more user-preferred interaction experience. Videos, appendixes and more are available at https://sites.google.com/view/vlm-handover/.
format Preprint
id arxiv_https___arxiv_org_abs_2409_17621
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human Handover
Liu, Jiangshan
Dong, Wenlong
Wang, Jiankun
Meng, Max Q. -H.
Robotics
Human-robot interaction (HRI) encompasses a wide range of collaborative tasks, with handover being one of the most fundamental. As robots become more integrated into human environments, the potential for service robots to assist in handing objects to humans is increasingly promising. In robot-to-human (R2H) handover, selecting the optimal grasp is crucial for success, as it requires avoiding interference with the humans preferred grasp region and minimizing intrusion into their workspace. Existing methods either inadequately consider geometric information or rely on data-driven approaches, which often struggle to generalize across diverse objects. To address these limitations, we propose a novel zero-shot system that combines semantic and geometric information to generate optimal handover grasps. Our method first identifies grasp regions using semantic knowledge from vision-language models (VLMs) and, by incorporating customized visual prompts, achieves finer granularity in region grounding. A grasp is then selected based on grasp distance and approach angle to maximize human ease and avoid interference. We validate our approach through ablation studies and real-world comparison experiments. Results demonstrate that our system improves handover success rates and provides a more user-preferred interaction experience. Videos, appendixes and more are available at https://sites.google.com/view/vlm-handover/.
title Leveraging Semantic and Geometric Information for Zero-Shot Robot-to-Human Handover
topic Robotics
url https://arxiv.org/abs/2409.17621