RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jinxuan, Jin, Shiyu, Lei, Yutian, Zhang, Yuqian, Zhang, Liangjun
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912110840119296
author Xu, Jinxuan
Jin, Shiyu
Lei, Yutian
Zhang, Yuqian
Zhang, Liangjun
author_facet Xu, Jinxuan
Jin, Shiyu
Lei, Yutian
Zhang, Yuqian
Zhang, Liangjun
contents Recent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses this limitation by investigating the potential of adopting the reasoning ability of LLMs for generating numerical predictions in robotics tasks, specifically for robotic grasping. We propose Reasoning Tuning, a novel method that integrates a reasoning phase before prediction during training, leveraging the extensive prior knowledge and advanced reasoning abilities of LLMs. This approach enables LLMs, notably with multi-modal capabilities, to generate accurate numerical outputs like grasp poses that are context-aware and adaptable through conversations. Additionally, we present the Reasoning Tuning VLM Grasp dataset, carefully curated to facilitate the adaptation of LLMs to robotic grasping. Extensive validation on both grasping datasets and real-world experiments underscores the adaptability of multi-modal LLMs for numerical prediction tasks in robotics. This not only expands their applicability but also bridges the gap between text-based planning and direct robot control, thereby maximizing the potential of LLMs in robotics.
format Preprint
id arxiv_https___arxiv_org_abs_2411_05212
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language Model
Xu, Jinxuan
Jin, Shiyu
Lei, Yutian
Zhang, Yuqian
Zhang, Liangjun
Robotics
Recent advances in Large Language Models (LLMs) have showcased their remarkable reasoning capabilities, making them influential across various fields. However, in robotics, their use has primarily been limited to manipulation planning tasks due to their inherent textual output. This paper addresses this limitation by investigating the potential of adopting the reasoning ability of LLMs for generating numerical predictions in robotics tasks, specifically for robotic grasping. We propose Reasoning Tuning, a novel method that integrates a reasoning phase before prediction during training, leveraging the extensive prior knowledge and advanced reasoning abilities of LLMs. This approach enables LLMs, notably with multi-modal capabilities, to generate accurate numerical outputs like grasp poses that are context-aware and adaptable through conversations. Additionally, we present the Reasoning Tuning VLM Grasp dataset, carefully curated to facilitate the adaptation of LLMs to robotic grasping. Extensive validation on both grasping datasets and real-world experiments underscores the adaptability of multi-modal LLMs for numerical prediction tasks in robotics. This not only expands their applicability but also bridges the gap between text-based planning and direct robot control, thereby maximizing the potential of LLMs in robotics.
title RT-Grasp: Reasoning Tuning Robotic Grasping via Multi-modal Large Language Model
topic Robotics
url https://arxiv.org/abs/2411.05212