QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Tong, Xinyang, Ding, Pengxiang, Fan, Yiguo, Wang, Donglin, Zhang, Wenjie, Cui, Can, Sun, Mingyang, Zhao, Han, Zhang, Hongyin, Dang, Yonghao, Huang, Siteng, Lyu, Shangke
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908380992372736
author Tong, Xinyang
Ding, Pengxiang
Fan, Yiguo
Wang, Donglin
Zhang, Wenjie
Cui, Can
Sun, Mingyang
Zhao, Han
Zhang, Hongyin
Dang, Yonghao
Huang, Siteng
Lyu, Shangke
author_facet Tong, Xinyang
Ding, Pengxiang
Fan, Yiguo
Wang, Donglin
Zhang, Wenjie
Cui, Can
Sun, Mingyang
Zhao, Han
Zhang, Hongyin
Dang, Yonghao
Huang, Siteng
Lyu, Shangke
contents This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUART-Online, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65%. Our project page is https://quart-online.github.io.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15576
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
Tong, Xinyang
Ding, Pengxiang
Fan, Yiguo
Wang, Donglin
Zhang, Wenjie
Cui, Can
Sun, Mingyang
Zhao, Han
Zhang, Hongyin
Dang, Yonghao
Huang, Siteng
Lyu, Shangke
Robotics
Computer Vision and Pattern Recognition
This paper addresses the inherent inference latency challenges associated with deploying multimodal large language models (MLLM) in quadruped vision-language-action (QUAR-VLA) tasks. Our investigation reveals that conventional parameter reduction techniques ultimately impair the performance of the language foundation model during the action instruction tuning phase, making them unsuitable for this purpose. We introduce a novel latency-free quadruped MLLM model, dubbed QUART-Online, designed to enhance inference efficiency without degrading the performance of the language foundation model. By incorporating Action Chunk Discretization (ACD), we compress the original action representation space, mapping continuous action values onto a smaller set of discrete representative vectors while preserving critical information. Subsequently, we fine-tune the MLLM to integrate vision, language, and compressed actions into a unified semantic space. Experimental results demonstrate that QUART-Online operates in tandem with the existing MLLM system, achieving real-time inference in sync with the underlying controller frequency, significantly boosting the success rate across various tasks by 65%. Our project page is https://quart-online.github.io.
title QUART-Online: Latency-Free Large Multimodal Language Model for Quadruped Robot Learning
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.15576