PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zheng, Yupeng, Li, Xiang, Gu, Songen, Zheng, Yuhang, Tian, Shuai, Li, Weize, Wang, Linbo, Fei, Senyu, Li, Pengfei, Gao, Yinfeng, Xing, Zebin, Chen, Yilun, Zhang, Qichao, Li, Haoran, Ding, Wenchao
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918465805221888
author Zheng, Yupeng
Li, Xiang
Gu, Songen
Zheng, Yuhang
Tian, Shuai
Li, Weize
Wang, Linbo
Fei, Senyu
Li, Pengfei
Gao, Yinfeng
Xing, Zebin
Chen, Yilun
Zhang, Qichao
Li, Haoran
Ding, Wenchao
author_facet Zheng, Yupeng
Li, Xiang
Gu, Songen
Zheng, Yuhang
Tian, Shuai
Li, Weize
Wang, Linbo
Fei, Senyu
Li, Pengfei
Gao, Yinfeng
Xing, Zebin
Chen, Yilun
Zhang, Qichao
Li, Haoran
Ding, Wenchao
contents Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we propose PokeVLA, a lightweight yet powerful foundation model for embodied manipulation that effectively infuses vision-language understanding into action learning. Our framework introduces a two-stage training paradigm: first, we pre-train a compact vision-language model (PokeVLM) on a curated multimodal dataset of 2.4M samples encompassing spatial grounding, affordance, and embodied reasoning tasks; second, we inject manipulation-relevant representations into the action space through multi-view goal-aware semantics learning, geometry alignment, and a novel action expert. Extensive experiments demonstrate state-of-the-art performance on the LIBERO-Plus benchmark and in real-world deployment, outperforming comparable baselines in success rate and robustness under diverse perturbations. To foster reproducibility and community progress, we will open-source our code, model weights, and the scripts for the curated pre-training dataset. Project page: https://getterupper.github.io/PokeVLA
format Preprint
id arxiv_https___arxiv_org_abs_2604_20834
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
Zheng, Yupeng
Li, Xiang
Gu, Songen
Zheng, Yuhang
Tian, Shuai
Li, Weize
Wang, Linbo
Fei, Senyu
Li, Pengfei
Gao, Yinfeng
Xing, Zebin
Chen, Yilun
Zhang, Qichao
Li, Haoran
Ding, Wenchao
Robotics
Recent advances in Vision-Language-Action (VLA) models have opened new avenues for robot manipulation, yet existing methods exhibit limited efficiency and a lack of high-level knowledge and spatial awareness. To address these challenges, we propose PokeVLA, a lightweight yet powerful foundation model for embodied manipulation that effectively infuses vision-language understanding into action learning. Our framework introduces a two-stage training paradigm: first, we pre-train a compact vision-language model (PokeVLM) on a curated multimodal dataset of 2.4M samples encompassing spatial grounding, affordance, and embodied reasoning tasks; second, we inject manipulation-relevant representations into the action space through multi-view goal-aware semantics learning, geometry alignment, and a novel action expert. Extensive experiments demonstrate state-of-the-art performance on the LIBERO-Plus benchmark and in real-world deployment, outperforming comparable baselines in success rate and robustness under diverse perturbations. To foster reproducibility and community progress, we will open-source our code, model weights, and the scripts for the curated pre-training dataset. Project page: https://getterupper.github.io/PokeVLA
title PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
topic Robotics
url https://arxiv.org/abs/2604.20834