Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Kechun, Zhu, Zhenjie, Chen, Anzhe, Zhao, Shuqi, Huang, Qing, Yang, Yifei, Lu, Haojian, Xiong, Rong, Tomizuka, Masayoshi, Wang, Yue
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918245676613632
author Xu, Kechun
Zhu, Zhenjie
Chen, Anzhe
Zhao, Shuqi
Huang, Qing
Yang, Yifei
Lu, Haojian
Xiong, Rong
Tomizuka, Masayoshi
Wang, Yue
author_facet Xu, Kechun
Zhu, Zhenjie
Chen, Anzhe
Zhao, Shuqi
Huang, Qing
Yang, Yifei
Lu, Haojian
Xiong, Rong
Tomizuka, Masayoshi
Wang, Yue
contents The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with external reasoning data helps, it requires experienced tuning and data-related overhead. Beyond such external dependencies, we identify an intrinsic cause within VLA datasets: modality imbalance, where language diversity is much lower than visual and action diversity. This imbalance biases the model toward visual shortcuts and language forgetting. To address this, we introduce BayesVLA, a Bayesian factorization that decomposes the policy into a visual-action prior, supporting seeing-to-act, and a language-conditioned likelihood, enabling prompt-to-specify. This inherently preserves generalization and promotes instruction following. We further incorporate pre- and post-contact phases to better leverage pre-trained foundation models. Information-theoretic analysis formally validates our effectiveness in mitigating shortcut learning. Extensive experiments show superior generalization to unseen instructions, objects, and environments compared to existing methods. Project page is available at: https://xukechun.github.io/papers/BayesVLA.
format Preprint
id arxiv_https___arxiv_org_abs_2512_11218
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
Xu, Kechun
Zhu, Zhenjie
Chen, Anzhe
Zhao, Shuqi
Huang, Qing
Yang, Yifei
Lu, Haojian
Xiong, Rong
Tomizuka, Masayoshi
Wang, Yue
Robotics
Computer Vision and Pattern Recognition
The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with external reasoning data helps, it requires experienced tuning and data-related overhead. Beyond such external dependencies, we identify an intrinsic cause within VLA datasets: modality imbalance, where language diversity is much lower than visual and action diversity. This imbalance biases the model toward visual shortcuts and language forgetting. To address this, we introduce BayesVLA, a Bayesian factorization that decomposes the policy into a visual-action prior, supporting seeing-to-act, and a language-conditioned likelihood, enabling prompt-to-specify. This inherently preserves generalization and promotes instruction following. We further incorporate pre- and post-contact phases to better leverage pre-trained foundation models. Information-theoretic analysis formally validates our effectiveness in mitigating shortcut learning. Extensive experiments show superior generalization to unseen instructions, objects, and environments compared to existing methods. Project page is available at: https://xukechun.github.io/papers/BayesVLA.
title Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
topic Robotics
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.11218