VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Weiqi, Zhang, Quande, Zhai, Ruifeng, Lin, Liang, Wang, Guangrun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912989612867584
author Li, Weiqi
Zhang, Quande
Zhai, Ruifeng
Lin, Liang
Wang, Guangrun
author_facet Li, Weiqi
Zhang, Quande
Zhai, Ruifeng
Lin, Liang
Wang, Guangrun
contents Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose a one-shot adaptation framework that recalibrates visual representations through lightweight, learnable updates. Our first method, Feature Token Modulation (FTM), applies a global affine transformation to visual tokens and improves Libero viewpoint accuracy from 48.5% to 87.1% with only 4K parameters. Building on this, Feature Linear Adaptation (FLA) introduces low-rank updates to the ViT encoder, achieving 90.8% success with 4.7M parameters -- matching LoRA-scale finetuning at far lower cost. Together, these results reveal substantial untapped robustness in pretrained VLA models and demonstrate that targeted, minimal visual adaptation is sufficient to restore viewpoint generalization.
format Preprint
id arxiv_https___arxiv_org_abs_2512_02902
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
Li, Weiqi
Zhang, Quande
Zhai, Ruifeng
Lin, Liang
Wang, Guangrun
Robotics
Artificial Intelligence
Machine Learning
Vision-language-action (VLA) models achieve strong in-distribution performance but degrade sharply under novel camera viewpoints and visual perturbations. We show that this brittleness primarily arises from misalignment in Spatial Modeling, rather than Physical Modeling. To address this, we propose a one-shot adaptation framework that recalibrates visual representations through lightweight, learnable updates. Our first method, Feature Token Modulation (FTM), applies a global affine transformation to visual tokens and improves Libero viewpoint accuracy from 48.5% to 87.1% with only 4K parameters. Building on this, Feature Linear Adaptation (FLA) introduces low-rank updates to the ViT encoder, achieving 90.8% success with 4.7M parameters -- matching LoRA-scale finetuning at far lower cost. Together, these results reveal substantial untapped robustness in pretrained VLA models and demonstrate that targeted, minimal visual adaptation is sufficient to restore viewpoint generalization.
title VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
topic Robotics
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2512.02902