Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Kechun, Zhu, Zhenjie, Chen, Anzhe, Zhao, Shuqi, Huang, Qing, Yang, Yifei, Lu, Haojian, Xiong, Rong, Tomizuka, Masayoshi, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Toward Embodiment Equivariant Vision-Language-Action Policy
by: Chen, Anzhe, et al.
Published: (2025)
by: Chen, Anzhe, et al.
Published: (2025)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024)
by: Xu, Kechun, et al.
Published: (2024)
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
by: Yang, Yifei, et al.
Published: (2026)
by: Yang, Yifei, et al.
Published: (2026)
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023)
by: Xu, Kechun, et al.
Published: (2023)
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
by: Xie, Liang, et al.
Published: (2022)
by: Xie, Liang, et al.
Published: (2022)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025)
by: Lou, Zhichen, et al.
Published: (2025)
Act to See, See to Act: Diffusion-Driven Perception-Action Interplay for Adaptive Policies
by: Wang, Jing, et al.
Published: (2025)
by: Wang, Jing, et al.
Published: (2025)
Revisit Mixture Models for Multi-Agent Simulation: Experimental Study within a Unified Framework
by: Lin, Longzhong, et al.
Published: (2025)
by: Lin, Longzhong, et al.
Published: (2025)
Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots
by: Zou, Ziqing, et al.
Published: (2026)
by: Zou, Ziqing, et al.
Published: (2026)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
Adaptive Linear Path Model-Based Diffusion
by: Shimizu, Yutaka, et al.
Published: (2026)
by: Shimizu, Yutaka, et al.
Published: (2026)
Leveraging Extrinsic Dexterity for Occluded Grasping on Grasp Constraining Walls
by: Kobashi, Keita, et al.
Published: (2025)
by: Kobashi, Keita, et al.
Published: (2025)
CLAW: Composable Language-Annotated Whole-body Motion Generation
by: Cao, Jianuo, et al.
Published: (2026)
by: Cao, Jianuo, et al.
Published: (2026)
Nonparametric Inverse Dynamic Models for Multimodal Interactive Robots
by: Haninger, Kevin, et al.
Published: (2019)
by: Haninger, Kevin, et al.
Published: (2019)
Fractional-order Modeling for Nonlinear Soft Actuators via Particle Swarm Optimization
by: Yang, Wu-Te, et al.
Published: (2025)
by: Yang, Wu-Te, et al.
Published: (2025)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
by: Jia, Xiaosong, et al.
Published: (2026)
by: Jia, Xiaosong, et al.
Published: (2026)
FDPP: Fine-tune Diffusion Policy with Human Preference
by: Chen, Yuxin, et al.
Published: (2025)
by: Chen, Yuxin, et al.
Published: (2025)
Bridging the Sim-to-Real Gap with Dynamic Compliance Tuning for Industrial Insertion
by: Zhang, Xiang, et al.
Published: (2023)
by: Zhang, Xiang, et al.
Published: (2023)
Language-Driven Policy Distillation for Cooperative Driving in Multi-Agent Reinforcement Learning
by: Liu, Jiaqi, et al.
Published: (2024)
by: Liu, Jiaqi, et al.
Published: (2024)
DexH2R: Task-oriented Dexterous Manipulation from Human to Robots
by: Zhao, Shuqi, et al.
Published: (2024)
by: Zhao, Shuqi, et al.
Published: (2024)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
Pretraining-finetuning Framework for Efficient Co-design: A Case Study on Quadruped Robot Parkour
by: Chen, Ci, et al.
Published: (2024)
by: Chen, Ci, et al.
Published: (2024)
Learning-Based Dynamics Modeling and Robust Control for Tendon-Driven Continuum Robots
by: Zou, Ziqing, et al.
Published: (2026)
by: Zou, Ziqing, et al.
Published: (2026)
Towards Generalizable and Interpretable Motion Prediction: A Deep Variational Bayes Approach
by: Lu, Juanwu, et al.
Published: (2024)
by: Lu, Juanwu, et al.
Published: (2024)
Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models
by: Ling, Yiran, et al.
Published: (2026)
by: Ling, Yiran, et al.
Published: (2026)
Underactuated Control of Multiple Soft Pneumatic Actuators via Stable Inversion
by: Yang, Wu-Te, et al.
Published: (2024)
by: Yang, Wu-Te, et al.
Published: (2024)
Optimized Design of a Soft Actuator Considering Force/Torque, Bendability, and Controllability via an Approximated Structure
by: Yang, Wu-Te, et al.
Published: (2023)
by: Yang, Wu-Te, et al.
Published: (2023)
Efficient Reinforcement Learning of Task Planners for Robotic Palletization through Iterative Action Masking Learning
by: Wu, Zheng, et al.
Published: (2024)
by: Wu, Zheng, et al.
Published: (2024)
Robust Model-Based In-Hand Manipulation with Integrated Real-Time Motion-Contact Planning and Tracking
by: Jiang, Yongpeng, et al.
Published: (2025)
by: Jiang, Yongpeng, et al.
Published: (2025)
Contact-Implicit Model Predictive Control for Dexterous In-hand Manipulation: A Long-Horizon and Robust Approach
by: Jiang, Yongpeng, et al.
Published: (2024)
by: Jiang, Yongpeng, et al.
Published: (2024)
Residual-MPPI: Online Policy Customization for Continuous Control
by: Wang, Pengcheng, et al.
Published: (2024)
by: Wang, Pengcheng, et al.
Published: (2024)
Robust In-Hand Manipulation with Extrinsic Contacts
by: Liang, Boyuan, et al.
Published: (2024)
by: Liang, Boyuan, et al.
Published: (2024)
Unified Manipulability and Compliance Analysis of Modular Soft-Rigid Hybrid Fingers
by: Zhou, Jianshu, et al.
Published: (2025)
by: Zhou, Jianshu, et al.
Published: (2025)
Frequency Domain Analysis of Nonlinear Series Elastic Actuator via Describing Function
by: Hirao, Motohiro, et al.
Published: (2023)
by: Hirao, Motohiro, et al.
Published: (2023)
P2 Explore: Efficient Exploration in Unknown Cluttered Environment with Floor Plan Prediction
by: Song, Kun, et al.
Published: (2024)
by: Song, Kun, et al.
Published: (2024)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
by: Wang, Zixuan, et al.
Published: (2026)
by: Wang, Zixuan, et al.
Published: (2026)
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning
by: Wang, Yixiao, et al.
Published: (2024)
by: Wang, Yixiao, et al.
Published: (2024)
PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models
by: Guo, Dingkun, et al.
Published: (2024)
by: Guo, Dingkun, et al.
Published: (2024)
Similar Items
-
Toward Embodiment Equivariant Vision-Language-Action Policy
by: Chen, Anzhe, et al.
Published: (2025) -
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024) -
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
by: Yang, Yifei, et al.
Published: (2026) -
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023) -
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
by: Xie, Liang, et al.
Published: (2022)