Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Grover, Shresth, Gopalkrishnan, Akshay, Ai, Bo, Christensen, Henrik I., Su, Hao, Li, Xuanlin
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866908543174574080
author Grover, Shresth
Gopalkrishnan, Akshay
Ai, Bo
Christensen, Henrik I.
Su, Hao
Li, Xuanlin
author_facet Grover, Shresth
Gopalkrishnan, Akshay
Ai, Bo
Christensen, Henrik I.
Su, Hao
Li, Xuanlin
contents Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, direct fine-tuning on robot data often disrupts these representations and limits generalization. We present a framework that better preserves pretrained features while adapting them for robot manipulation. Our approach introduces three components: (i) a dual-encoder design with one frozen vision encoder to retain pretrained features and another trainable for task adaptation, (ii) a string-based action tokenizer that casts continuous actions into character sequences aligned with the model's pretraining domain, and (iii) a co-training strategy that combines robot demonstrations with vision-language datasets emphasizing spatial reasoning and affordances. Evaluations in simulation and on real robots show that our method improves robustness to visual perturbations, generalization to novel instructions and environments, and overall task success compared to baselines.
format Preprint
id arxiv_https___arxiv_org_abs_2509_11417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
Grover, Shresth
Gopalkrishnan, Akshay
Ai, Bo
Christensen, Henrik I.
Su, Hao
Li, Xuanlin
Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, direct fine-tuning on robot data often disrupts these representations and limits generalization. We present a framework that better preserves pretrained features while adapting them for robot manipulation. Our approach introduces three components: (i) a dual-encoder design with one frozen vision encoder to retain pretrained features and another trainable for task adaptation, (ii) a string-based action tokenizer that casts continuous actions into character sequences aligned with the model's pretraining domain, and (iii) a co-training strategy that combines robot demonstrations with vision-language datasets emphasizing spatial reasoning and affordances. Evaluations in simulation and on real robots show that our method improves robustness to visual perturbations, generalization to novel instructions and environments, and overall task success compared to baselines.
title Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
topic Robotics
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2509.11417