NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Ziyue, Wu, Shangyang, Zhao, Shuai, Zhao, Zhiqiu, Li, Shengjie, Wang, Yi, Li, Fang, Luo, Haoran
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915850494148608
author Zhu, Ziyue
Wu, Shangyang
Zhao, Shuai
Zhao, Zhiqiu
Li, Shengjie
Wang, Yi
Li, Fang
Luo, Haoran
author_facet Zhu, Ziyue
Wu, Shangyang
Zhao, Shuai
Zhao, Zhiqiu
Li, Shengjie
Wang, Yi
Li, Fang
Luo, Haoran
contents Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable primitives, reducing reliance on large-scale data and complex architectures, and enabling exploration beyond demonstrations. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework via online reinforcement learning (RL). It introduces a symbolic encoder to embedding vision and language features and extract structured primitives, utilizes a symbolic solver for data-efficient action sequencing, and leverages online RL to optimize generation via expansive exploration. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability, high data efficiency and expanded exploration space. Our code is available.
format Preprint
id arxiv_https___arxiv_org_abs_2603_09542
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
Zhu, Ziyue
Wu, Shangyang
Zhao, Shuai
Zhao, Zhiqiu
Li, Shengjie
Wang, Yi
Li, Fang
Luo, Haoran
Robotics
Vision-Language-Action (VLA) models are formulated to ground instructions in visual context and generate action sequences for robotic manipulation. Despite recent progress, VLA models still face challenges in learning related and reusable primitives, reducing reliance on large-scale data and complex architectures, and enabling exploration beyond demonstrations. To address these challenges, we propose a novel Neuro-Symbolic Vision-Language-Action (NS-VLA) framework via online reinforcement learning (RL). It introduces a symbolic encoder to embedding vision and language features and extract structured primitives, utilizes a symbolic solver for data-efficient action sequencing, and leverages online RL to optimize generation via expansive exploration. Experiments on robotic manipulation benchmarks demonstrate that NS-VLA outperforms previous methods in both one-shot training and data-perturbed settings, while simultaneously exhibiting superior zero-shot generalizability, high data efficiency and expanded exploration space. Our code is available.
title NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
topic Robotics
url https://arxiv.org/abs/2603.09542