SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Shukor, Mustafa, Aubakirova, Dana, Capuano, Francesco, Kooijmans, Pepijn, Palma, Steven, Zouitine, Adil, Aractingi, Michel, Pascal, Caroline, Russi, Martino, Marafioti, Andres, Alibert, Simon, Cord, Matthieu, Wolf, Thomas, Cadene, Remi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LeRobot: An Open-Source Library for End-to-End Robot Learning
by: Cadene, Remi, et al.
Published: (2026)
by: Cadene, Remi, et al.
Published: (2026)
Robot Learning: A Tutorial
by: Capuano, Francesco, et al.
Published: (2025)
by: Capuano, Francesco, et al.
Published: (2025)
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024)
by: Shukor, Mustafa, et al.
Published: (2024)
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)
by: Vallaeys, Théophane, et al.
Published: (2024)
Beyond Task Performance: Evaluating and Reducing the Flaws of Large Multimodal Models with In-Context Learning
by: Shukor, Mustafa, et al.
Published: (2023)
by: Shukor, Mustafa, et al.
Published: (2023)
Solving robust MDPs as a sequence of static RL problems
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Analyzing Finetuning Representation Shift for Multimodal LLMs Steering
by: Khayatan, Pegah, et al.
Published: (2025)
by: Khayatan, Pegah, et al.
Published: (2025)
A Concept-Based Explainability Framework for Large Multimodal Models
by: Parekh, Jayneel, et al.
Published: (2024)
by: Parekh, Jayneel, et al.
Published: (2024)
DiffCut: Catalyzing Zero-Shot Semantic Segmentation with Diffusion Features and Recursive Normalized Cut
by: Couairon, Paul, et al.
Published: (2024)
by: Couairon, Paul, et al.
Published: (2024)
What Makes Multimodal In-Context Learning Work?
by: Baldassini, Folco Bertini, et al.
Published: (2024)
by: Baldassini, Folco Bertini, et al.
Published: (2024)
Learning to Steer: Input-dependent Steering for Multimodal LLMs
by: Parekh, Jayneel, et al.
Published: (2025)
by: Parekh, Jayneel, et al.
Published: (2025)
When Prompts Override Vision: Prompt-Induced Hallucinations in LVLMs
by: Khayatan, Pegah, et al.
Published: (2026)
by: Khayatan, Pegah, et al.
Published: (2026)
FreeSeg-Diff: Training-Free Open-Vocabulary Segmentation with Diffusion Models
by: Corradini, Barbara Toniella, et al.
Published: (2024)
by: Corradini, Barbara Toniella, et al.
Published: (2024)
RRLS : Robust Reinforcement Learning Suite
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
Time-Constrained Robust MDPs
by: Zouitine, Adil, et al.
Published: (2024)
by: Zouitine, Adil, et al.
Published: (2024)
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
by: Allal, Loubna Ben, et al.
Published: (2025)
by: Allal, Loubna Ben, et al.
Published: (2025)
SmolVLM: Redefining small and efficient multimodal models
by: Marafioti, Andrés, et al.
Published: (2025)
by: Marafioti, Andrés, et al.
Published: (2025)
Scaling Laws for Native Multimodal Models
by: Shukor, Mustafa, et al.
Published: (2025)
by: Shukor, Mustafa, et al.
Published: (2025)
Back to the Baseline: Examining Baseline Effects on Explainability Metrics
by: Picard, Agustin Martin, et al.
Published: (2025)
by: Picard, Agustin Martin, et al.
Published: (2025)
SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion
by: Nassar, Ahmed, et al.
Published: (2025)
by: Nassar, Ahmed, et al.
Published: (2025)
INFLUENCE OF WATER MINERALIZATION ON ZOOPLANKTON PRODUCTIVITY IN RESERVOIRS OF AKMOLA REGION
by: Aubakirova, Gulzhan, et al.
Published: (2020)
by: Aubakirova, Gulzhan, et al.
Published: (2020)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
by: Chopra, Samarth, et al.
Published: (2025)
by: Chopra, Samarth, et al.
Published: (2025)
FIN DE AÑO SIN FIN
by: Roberto Marafioti
Published: (2005)
by: Roberto Marafioti
Published: (2005)
Smol-GS: Compact Representations for Abstract 3D Gaussian Splatting
by: Wang, Haishan, et al.
Published: (2025)
by: Wang, Haishan, et al.
Published: (2025)
The Rise and Fall of the People's Parties
by: Corduwener, Pepijn
Published: (2023)
by: Corduwener, Pepijn
Published: (2023)
SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization
by: Vallaeys, Théophane, et al.
Published: (2025)
by: Vallaeys, Théophane, et al.
Published: (2025)
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
by: Messaoud, Kaouther, et al.
Published: (2025)
by: Messaoud, Kaouther, et al.
Published: (2025)
SmolRGPT: Efficient Spatial Reasoning for Warehouse Environments with 600M Parameters
by: Traore, Abdarahmane, et al.
Published: (2025)
by: Traore, Abdarahmane, et al.
Published: (2025)
A dimensão comunicacional como recorte metodológico para o estudo das migrações
by: Pedro Russi
Published: (2014)
by: Pedro Russi
Published: (2014)
Tamiriki, pata yotono kwama: a reconstrução de uma casa, a valorização de uma cultura e o protagonismo dos ameríndios Kaxuyana às margens do rio Cachorro (Oriximiná-Pa)
by: Adriana Russi
Published: (2017)
by: Adriana Russi
Published: (2017)
Los pasivos ambientales
by: Daniela Russi
Published: (2002)
by: Daniela Russi
Published: (2002)
La producción del espacio urbano en la vida de un grupo de migrantes peruanas, empleadas domésticas en Brasilia
by: Pedro Russi
Published: (2013)
by: Pedro Russi
Published: (2013)
ARQUITETURA DO ESPAÇO ESCOLAR, ADEQUAÇÃO DA EDIFICAÇÃO AOS PARÂMETROS AMBIENTAIS: ESTUDO DE CASO CTISM-COLÉGIO TÉCNICO INDUSTRIAL DE SANTA MARIA
by: Madalena Russi
Published: (2014)
by: Madalena Russi
Published: (2014)
Epígrafes. Imaginaciones deseables y otras epistemes
by: Pedro Russi
Published: (2020)
by: Pedro Russi
Published: (2020)
PROBLEMÁTICAS CONCERNIENTES A LA RELACIÓN COMUNICACIÓN-MIGRACIÓN
by: Pedro Russi
Published: (2016)
by: Pedro Russi
Published: (2016)
Coleções etnográficas, povos indígenas e práticas de representação: as mudanças nos processos museais com as experiências colaborativas
by: Adriana Russi
Published: (2018)
by: Adriana Russi
Published: (2018)
Regímenes latentes de error en el aprendizaje de la concordancia plural en ELE
by: Pablo Ezequiel Marafioti
Published: (2024)
by: Pablo Ezequiel Marafioti
Published: (2024)
ANÁLISIS DE LA EVOLUCIÓN DE ERRORES DE CONCORDANCIA EN CUATRO APRENDIENTES ITALIANOS DE ELE USANDO REDES COMPLEJAS*
by: Pablo Ezequiel Marafioti
Published: (2021)
by: Pablo Ezequiel Marafioti
Published: (2021)
Análisis de tiempos hasta que se produce un error de concordancia en cuatro estudiantes italianos de ELE
by: Pablo Ezequiel Marafioti
Published: (2022)
by: Pablo Ezequiel Marafioti
Published: (2022)
Similar Items
-
LeRobot: An Open-Source Library for End-to-End Robot Learning
by: Cadene, Remi, et al.
Published: (2026) -
Robot Learning: A Tutorial
by: Capuano, Francesco, et al.
Published: (2025) -
Implicit Multimodal Alignment: On the Generalization of Frozen LLMs to Multimodal Inputs
by: Shukor, Mustafa, et al.
Published: (2024) -
Skipping Computations in Multimodal LLMs
by: Shukor, Mustafa, et al.
Published: (2024) -
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs
by: Vallaeys, Théophane, et al.
Published: (2024)