Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866918245676613632 |
|---|---|
| author | Xu, Kechun Zhu, Zhenjie Chen, Anzhe Zhao, Shuqi Huang, Qing Yang, Yifei Lu, Haojian Xiong, Rong Tomizuka, Masayoshi Wang, Yue |
| author_facet | Xu, Kechun Zhu, Zhenjie Chen, Anzhe Zhao, Shuqi Huang, Qing Yang, Yifei Lu, Haojian Xiong, Rong Tomizuka, Masayoshi Wang, Yue |
| contents | The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with external reasoning data helps, it requires experienced tuning and data-related overhead. Beyond such external dependencies, we identify an intrinsic cause within VLA datasets: modality imbalance, where language diversity is much lower than visual and action diversity. This imbalance biases the model toward visual shortcuts and language forgetting. To address this, we introduce BayesVLA, a Bayesian factorization that decomposes the policy into a visual-action prior, supporting seeing-to-act, and a language-conditioned likelihood, enabling prompt-to-specify. This inherently preserves generalization and promotes instruction following. We further incorporate pre- and post-contact phases to better leverage pre-trained foundation models. Information-theoretic analysis formally validates our effectiveness in mitigating shortcut learning. Extensive experiments show superior generalization to unseen instructions, objects, and environments compared to existing methods. Project page is available at: https://xukechun.github.io/papers/BayesVLA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2512_11218 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy Xu, Kechun Zhu, Zhenjie Chen, Anzhe Zhao, Shuqi Huang, Qing Yang, Yifei Lu, Haojian Xiong, Rong Tomizuka, Masayoshi Wang, Yue Robotics Computer Vision and Pattern Recognition The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with external reasoning data helps, it requires experienced tuning and data-related overhead. Beyond such external dependencies, we identify an intrinsic cause within VLA datasets: modality imbalance, where language diversity is much lower than visual and action diversity. This imbalance biases the model toward visual shortcuts and language forgetting. To address this, we introduce BayesVLA, a Bayesian factorization that decomposes the policy into a visual-action prior, supporting seeing-to-act, and a language-conditioned likelihood, enabling prompt-to-specify. This inherently preserves generalization and promotes instruction following. We further incorporate pre- and post-contact phases to better leverage pre-trained foundation models. Information-theoretic analysis formally validates our effectiveness in mitigating shortcut learning. Extensive experiments show superior generalization to unseen instructions, objects, and environments compared to existing methods. Project page is available at: https://xukechun.github.io/papers/BayesVLA. |
| title | Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy |
| topic | Robotics Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2512.11218 |