SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shukor, Mustafa, Aubakirova, Dana, Capuano, Francesco, Kooijmans, Pepijn, Palma, Steven, Zouitine, Adil, Aractingi, Michel, Pascal, Caroline, Russi, Martino, Marafioti, Andres, Alibert, Simon, Cord, Matthieu, Wolf, Thomas, Cadene, Remi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915318365945856
author Shukor, Mustafa
Aubakirova, Dana
Capuano, Francesco
Kooijmans, Pepijn
Palma, Steven
Zouitine, Adil
Aractingi, Michel
Pascal, Caroline
Russi, Martino
Marafioti, Andres
Alibert, Simon
Cord, Matthieu
Wolf, Thomas
Cadene, Remi
author_facet Shukor, Mustafa
Aubakirova, Dana
Capuano, Francesco
Kooijmans, Pepijn
Palma, Steven
Zouitine, Adil
Aractingi, Michel
Pascal, Caroline
Russi, Martino
Marafioti, Andres
Alibert, Simon
Cord, Matthieu
Wolf, Thomas
Cadene, Remi
contents Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive--often with billions of parameters--leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10x larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic benchmarks and release all code, pretrained models, and training data.
format Preprint
id arxiv_https___arxiv_org_abs_2506_01844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Shukor, Mustafa
Aubakirova, Dana
Capuano, Francesco
Kooijmans, Pepijn
Palma, Steven
Zouitine, Adil
Aractingi, Michel
Pascal, Caroline
Russi, Martino
Marafioti, Andres
Alibert, Simon
Cord, Matthieu
Wolf, Thomas
Cadene, Remi
Machine Learning
Robotics
Vision-language models (VLMs) pretrained on large-scale multimodal datasets encode rich visual and linguistic knowledge, making them a strong foundation for robotics. Rather than training robotic policies from scratch, recent approaches adapt VLMs into vision-language-action (VLA) models that enable natural language-driven perception and control. However, existing VLAs are typically massive--often with billions of parameters--leading to high training costs and limited real-world deployability. Moreover, they rely on academic and industrial datasets, overlooking the growing availability of community-collected data from affordable robotic platforms. In this work, we present SmolVLA, a small, efficient, and community-driven VLA that drastically reduces both training and inference costs, while retaining competitive performance. SmolVLA is designed to be trained on a single GPU and deployed on consumer-grade GPUs or even CPUs. To further improve responsiveness, we introduce an asynchronous inference stack decoupling perception and action prediction from action execution, allowing higher control rates with chunked action generation. Despite its compact size, SmolVLA achieves performance comparable to VLAs that are 10x larger. We evaluate SmolVLA on a range of both simulated as well as real-world robotic benchmarks and release all code, pretrained models, and training data.
title SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
topic Machine Learning
Robotics
url https://arxiv.org/abs/2506.01844