Green-VLA: Staged Vision-Language-Action Model for Generalist Robots

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Apanasevich, I., Artemyev, M., Babakyan, R., Fedotova, P., Grankin, D., Kupryashin, E., Misailidi, A., Nerus, D., Nutalapati, A., Sidorov, G., Efremov, I., Gerasyov, M., Pikurov, D., Senchenko, Y., Davidenko, S., Kulikov, D., Sultankin, M., Askarbek, K., Shamanin, O., Statovoy, D., Zalyaev, E., Zorin, I., Letkin, A., Rusakov, E., Silchenko, A., Vorobyov, V., Sobolnikov, S., Postnikov, A.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915841488977920
author Apanasevich, I.
Artemyev, M.
Babakyan, R.
Fedotova, P.
Grankin, D.
Kupryashin, E.
Misailidi, A.
Nerus, D.
Nutalapati, A.
Sidorov, G.
Efremov, I.
Gerasyov, M.
Pikurov, D.
Senchenko, Y.
Davidenko, S.
Kulikov, D.
Sultankin, M.
Askarbek, K.
Shamanin, O.
Statovoy, D.
Zalyaev, E.
Zorin, I.
Letkin, A.
Rusakov, E.
Silchenko, A.
Vorobyov, V.
Sobolnikov, S.
Postnikov, A.
author_facet Apanasevich, I.
Artemyev, M.
Babakyan, R.
Fedotova, P.
Grankin, D.
Kupryashin, E.
Misailidi, A.
Nerus, D.
Nutalapati, A.
Sidorov, G.
Efremov, I.
Gerasyov, M.
Pikurov, D.
Senchenko, Y.
Davidenko, S.
Kulikov, D.
Sultankin, M.
Askarbek, K.
Shamanin, O.
Statovoy, D.
Zalyaev, E.
Zorin, I.
Letkin, A.
Rusakov, E.
Silchenko, A.
Vorobyov, V.
Sobolnikov, S.
Postnikov, A.
contents We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.
format Preprint
id arxiv_https___arxiv_org_abs_2602_00919
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
Apanasevich, I.
Artemyev, M.
Babakyan, R.
Fedotova, P.
Grankin, D.
Kupryashin, E.
Misailidi, A.
Nerus, D.
Nutalapati, A.
Sidorov, G.
Efremov, I.
Gerasyov, M.
Pikurov, D.
Senchenko, Y.
Davidenko, S.
Kulikov, D.
Sultankin, M.
Askarbek, K.
Shamanin, O.
Statovoy, D.
Zalyaev, E.
Zorin, I.
Letkin, A.
Rusakov, E.
Silchenko, A.
Vorobyov, V.
Sobolnikov, S.
Postnikov, A.
Robotics
We introduce Green-VLA, a staged Vision-Language-Action (VLA) framework for real-world deployment on the Green humanoid robot while maintaining generalization across diverse embodiments. Green-VLA follows a five stage curriculum: (L0) foundational VLMs, (L1) multimodal grounding, (R0) multi-embodiment pretraining, (R1) embodiment-specific adaptation, and (R2) reinforcement-learning (RL) policy alignment. We couple a scalable data-processing pipeline (3,000 hours of demonstrations) with temporal alignment and quality filtering, and use a unified, embodiment-aware action interface enabling a single policy to control humanoids, mobile manipulators, and fixed-base arms. At inference, the VLA controller is enhanced with episode-progress prediction, out-of-distribution detection, and joint-prediction-based guidance to improve safety and precise target selection. Experiments on Simpler BRIDGE WidowX and CALVIN ABC-D, as well as real-robot evaluations, show strong generalization and performance gains from RL alignment in success rate, robustness, and long-horizon efficiency.
title Green-VLA: Staged Vision-Language-Action Model for Generalist Robots
topic Robotics
url https://arxiv.org/abs/2602.00919