QUAR-VLA: Vision-Language-Action Model for Quadruped Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ding, Pengxiang, Zhao, Han, Zhang, Wenjie, Song, Wenxuan, Zhang, Min, Huang, Siteng, Yang, Ningxi, Wang, Donglin
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910812629630976
author Ding, Pengxiang
Zhao, Han
Zhang, Wenjie
Song, Wenxuan
Zhang, Min
Huang, Siteng
Yang, Ningxi
Wang, Donglin
author_facet Ding, Pengxiang
Zhao, Han
Zhang, Wenjie
Song, Wenxuan
Zhang, Min
Huang, Siteng
Yang, Ningxi
Wang, Donglin
contents The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentalize perception, planning, and decision-making, simplifying system design but limiting the synergy between different information streams. This compartmentalization poses challenges in achieving seamless autonomous reasoning, decision-making, and action execution. To address these limitations, a novel paradigm, named Vision-Language-Action tasks for QUAdruped Robots (QUAR-VLA), has been introduced in this paper. This approach tightly integrates visual information and instructions to generate executable actions, effectively merging perception, planning, and decision-making. The central idea is to elevate the overall intelligence of the robot. Within this framework, a notable challenge lies in aligning fine-grained instructions with visual perception information. This emphasizes the complexity involved in ensuring that the robot accurately interprets and acts upon detailed instructions in harmony with its visual observations. Consequently, we propose QUAdruped Robotic Transformer (QUART), a family of VLA models to integrate visual information and instructions from diverse modalities as input and generates executable actions for real-world robots and present QUAdruped Robot Dataset (QUARD), a large-scale multi-task dataset including navigation, complex terrain locomotion, and whole-body manipulation tasks for training QUART models. Our extensive evaluation (4000 evaluation trials) shows that our approach leads to performant robotic policies and enables QUART to obtain a range of emergent capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2312_14457
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
Ding, Pengxiang
Zhao, Han
Zhang, Wenjie
Song, Wenxuan
Zhang, Min
Huang, Siteng
Yang, Ningxi
Wang, Donglin
Robotics
Computer Vision and Pattern Recognition
The important manifestation of robot intelligence is the ability to naturally interact and autonomously make decisions. Traditional approaches to robot control often compartmentalize perception, planning, and decision-making, simplifying system design but limiting the synergy between different information streams. This compartmentalization poses challenges in achieving seamless autonomous reasoning, decision-making, and action execution. To address these limitations, a novel paradigm, named Vision-Language-Action tasks for QUAdruped Robots (QUAR-VLA), has been introduced in this paper. This approach tightly integrates visual information and instructions to generate executable actions, effectively merging perception, planning, and decision-making. The central idea is to elevate the overall intelligence of the robot. Within this framework, a notable challenge lies in aligning fine-grained instructions with visual perception information. This emphasizes the complexity involved in ensuring that the robot accurately interprets and acts upon detailed instructions in harmony with its visual observations. Consequently, we propose QUAdruped Robotic Transformer (QUART), a family of VLA models to integrate visual information and instructions from diverse modalities as input and generates executable actions for real-world robots and present QUAdruped Robot Dataset (QUARD), a large-scale multi-task dataset including navigation, complex terrain locomotion, and whole-body manipulation tasks for training QUART models. Our extensive evaluation (4000 evaluation trials) shows that our approach leads to performant robotic policies and enables QUART to obtain a range of emergent capabilities.
title QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.14457