VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tang, Zuojin, Hu, Bin, Zhao, Chenyang, Ma, De, Pan, Gang, Liu, Bin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750309945344
author Tang, Zuojin
Hu, Bin
Zhao, Chenyang
Ma, De
Pan, Gang
Liu, Bin
author_facet Tang, Zuojin
Hu, Bin
Zhao, Chenyang
Ma, De
Pan, Gang
Liu, Bin
contents Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MIMO) scenarios, where parallel task execution is required. In MISO architectures, tasks compete for a shared output channel, creating mutual exclusion effects that cause unbalanced optimization and degraded performance. To address this gap, we introduce MIMO-VLA (VLASCD), a unified training framework that enables concurrent multi-task outputs, exemplified by simultaneous dialogue generation and decision-making. Inspired by human cognition, MIMO-VLA eliminates interference between tasks and supports efficient parallel processing. Experiments on the CARLA autonomous driving platform demonstrate that MIMO-VLA substantially outperforms state-of-the-art MISO-based LLMs, reinforcement learning models, and VLAs in MIMO settings, establishing a new direction for multimodal and multitask learning.
format Preprint
id arxiv_https___arxiv_org_abs_2410_15885
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
Tang, Zuojin
Hu, Bin
Zhao, Chenyang
Ma, De
Pan, Gang
Liu, Bin
Artificial Intelligence
Recent large pretrained models such as LLMs (e.g., GPT series) and VLAs (e.g., OpenVLA) have achieved notable progress on multimodal tasks, yet they are built upon a multi-input single-output (MISO) paradigm. We show that this paradigm fundamentally limits performance in multi-input multi-output (MIMO) scenarios, where parallel task execution is required. In MISO architectures, tasks compete for a shared output channel, creating mutual exclusion effects that cause unbalanced optimization and degraded performance. To address this gap, we introduce MIMO-VLA (VLASCD), a unified training framework that enables concurrent multi-task outputs, exemplified by simultaneous dialogue generation and decision-making. Inspired by human cognition, MIMO-VLA eliminates interference between tasks and supports efficient parallel processing. Experiments on the CARLA autonomous driving platform demonstrate that MIMO-VLA substantially outperforms state-of-the-art MISO-based LLMs, reinforcement learning models, and VLAs in MIMO settings, establishing a new direction for multimodal and multitask learning.
title VLASCD: A Visual Language Action Model for Simultaneous Chatting and Decision Making
topic Artificial Intelligence
url https://arxiv.org/abs/2410.15885