VLANeXt: Recipes for Building Strong VLA Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Xiao-Ming, Fan, Bin, Liao, Kang, Jiang, Jian-Jian, Yang, Runze, Luo, Yihang, Wu, Zhonghua, Zheng, Wei-Shi, Loy, Chen Change
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913147002028032
author Wu, Xiao-Ming
Fan, Bin
Liao, Kang
Jiang, Jian-Jian
Yang, Runze
Luo, Yihang
Wu, Zhonghua
Zheng, Wei-Shi
Loy, Chen Change
author_facet Wu, Xiao-Ming
Fan, Bin
Liao, Kang
Jiang, Jian-Jian
Yang, Runze
Luo, Yihang
Wu, Zhonghua
Zheng, Wei-Shi
Loy, Chen Change
contents Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT-2, which is the origin of VLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. It outperforms the state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks and demonstrates strong performance in real-world experiments. We release a unified and easy-to-use codebase to reproduce our findings, explore the design space, and develop new VLA variants on top of a shared foundation. The codebase is available at https://github.com/DravenALG/VLANeXt.
format Preprint
id arxiv_https___arxiv_org_abs_2602_18532
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VLANeXt: Recipes for Building Strong VLA Models
Wu, Xiao-Ming
Fan, Bin
Liao, Kang
Jiang, Jian-Jian
Yang, Runze
Luo, Yihang
Wu, Zhonghua
Zheng, Wei-Shi
Loy, Chen Change
Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
Following the rise of large foundation models, Vision-Language-Action models (VLAs) emerged, leveraging strong visual and language understanding from Vision-Language Models for general-purpose policy learning. Yet, the current VLA landscape remains fragmented and exploratory. Although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. To bring structure to this evolving space, we reexamine the VLA design space under a unified framework and evaluation setup. Starting from a simple VLA baseline similar to RT-2, which is the origin of VLA, we systematically dissect design choices along three dimensions: foundational components, perception essentials, and action modelling perspectives. From this study, we distill 12 key findings that together form a practical recipe for building strong VLA models. The outcome of this exploration is a simple yet effective model, VLANeXt. It outperforms the state-of-the-art methods on the LIBERO and LIBERO-plus benchmarks and demonstrates strong performance in real-world experiments. We release a unified and easy-to-use codebase to reproduce our findings, explore the design space, and develop new VLA variants on top of a shared foundation. The codebase is available at https://github.com/DravenALG/VLANeXt.
title VLANeXt: Recipes for Building Strong VLA Models
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Robotics
url https://arxiv.org/abs/2602.18532