Toward Embodiment Equivariant Vision-Language-Action Policy
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Anzhe, Yang, Yifei, Zhu, Zhenjie, Xu, Kechun, Zhou, Zhongxiang, Xiong, Rong, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025)
by: Lou, Zhichen, et al.
Published: (2025)
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023)
by: Xu, Kechun, et al.
Published: (2023)
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
by: Yang, Yifei, et al.
Published: (2026)
by: Yang, Yifei, et al.
Published: (2026)
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024)
by: Xu, Kechun, et al.
Published: (2024)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
CNSv2: Probabilistic Correspondence Encoded Neural Image Servo
by: Chen, Anzhe, et al.
Published: (2025)
by: Chen, Anzhe, et al.
Published: (2025)
Embodiment Transfer Learning for Vision-Language-Action Models
by: Li, Chengmeng, et al.
Published: (2025)
by: Li, Chengmeng, et al.
Published: (2025)
Learning A Simulation-based Visual Policy for Real-world Peg In Unseen Holes
by: Xie, Liang, et al.
Published: (2022)
by: Xie, Liang, et al.
Published: (2022)
Disambiguate Gripper State in Grasp-Based Tasks: Pseudo-Tactile as Feedback Enables Pure Simulation Learning
by: Yang, Yifei, et al.
Published: (2025)
by: Yang, Yifei, et al.
Published: (2025)
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
by: Wang, Qiuyue, et al.
Published: (2026)
by: Wang, Qiuyue, et al.
Published: (2026)
TOP: Time Optimization Policy for Stable and Accurate Standing Manipulation with Humanoid Robots
by: Chen, Zhenghan, et al.
Published: (2025)
by: Chen, Zhenghan, et al.
Published: (2025)
One-Policy-Fits-All: Geometry-Aware Action Latents for Cross-Embodiment Manipulation
by: Mu, Juncheng, et al.
Published: (2026)
by: Mu, Juncheng, et al.
Published: (2026)
ACoT-VLA: Action Chain-of-Thought for Vision-Language-Action Models
by: Zhong, Linqing, et al.
Published: (2026)
by: Zhong, Linqing, et al.
Published: (2026)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
ELAN4D: Embodiment-Centric 4D Supervision for Vision-Language-Action Models via Plug-and-Play Adaptation
by: He, Zeyuan, et al.
Published: (2026)
by: He, Zeyuan, et al.
Published: (2026)
Bridging Embodiment Gaps: Deploying Vision-Language-Action Models on Soft Robots
by: Su, Haochen, et al.
Published: (2025)
by: Su, Haochen, et al.
Published: (2025)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
by: Yang, Rushuai, et al.
Published: (2026)
by: Yang, Rushuai, et al.
Published: (2026)
Revisit Mixture Models for Multi-Agent Simulation: Experimental Study within a Unified Framework
by: Lin, Longzhong, et al.
Published: (2025)
by: Lin, Longzhong, et al.
Published: (2025)
UMI-on-Air: Embodiment-Aware Guidance for Embodiment-Agnostic Visuomotor Policies
by: Gupta, Harsh, et al.
Published: (2025)
by: Gupta, Harsh, et al.
Published: (2025)
Bridging the Embodiment Gap: Disentangled Cross-Embodiment Video Editing
by: Li, Zhiyuan, et al.
Published: (2026)
by: Li, Zhiyuan, et al.
Published: (2026)
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
by: Zheng, Jinliang, et al.
Published: (2025)
by: Zheng, Jinliang, et al.
Published: (2025)
Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models
by: Zhang, Zhilong, et al.
Published: (2026)
by: Zhang, Zhilong, et al.
Published: (2026)
Latent Action Diffusion for Cross-Embodiment Manipulation
by: Bauer, Erik, et al.
Published: (2025)
by: Bauer, Erik, et al.
Published: (2025)
Dita: Scaling Diffusion Transformer for Generalist Vision-Language-Action Policy
by: Hou, Zhi, et al.
Published: (2025)
by: Hou, Zhi, et al.
Published: (2025)
FLOWER: Democratizing Generalist Robot Policies with Efficient Vision-Language-Action Flow Policies
by: Reuss, Moritz, et al.
Published: (2025)
by: Reuss, Moritz, et al.
Published: (2025)
Mirage: Cross-Embodiment Zero-Shot Policy Transfer with Cross-Painting
by: Chen, Lawrence Yunliang, et al.
Published: (2024)
by: Chen, Lawrence Yunliang, et al.
Published: (2024)
NS-VLA: Towards Neuro-Symbolic Vision-Language-Action Models
by: Zhu, Ziyue, et al.
Published: (2026)
by: Zhu, Ziyue, et al.
Published: (2026)
ActionFlow: A Pipelined Action Acceleration for Vision Language Models on Edge
by: Dai, Yuntao, et al.
Published: (2025)
by: Dai, Yuntao, et al.
Published: (2025)
HarvestFlex: Strawberry Harvesting via Vision-Language-Action Policy Adaptation in the Wild
by: Zhao, Ziyang, et al.
Published: (2026)
by: Zhao, Ziyang, et al.
Published: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
MOTIF: Learning Action Motifs for Few-shot Cross-Embodiment Transfer
by: Zhi, Heng, et al.
Published: (2026)
by: Zhi, Heng, et al.
Published: (2026)
RING#: PR-by-PE Global Localization with Roto-translation Equivariant Gram Learning
by: Lu, Sha, et al.
Published: (2024)
by: Lu, Sha, et al.
Published: (2024)
Embodiment-Agnostic Action Planning via Object-Part Scene Flow
by: Tang, Weiliang, et al.
Published: (2024)
by: Tang, Weiliang, et al.
Published: (2024)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
by: Zuo, Kuangji, et al.
Published: (2026)
by: Zuo, Kuangji, et al.
Published: (2026)
Embodied Interpretability: Linking Causal Understanding to Generalization in Vision-Language-Action Models
by: Zhang, Hanxin, et al.
Published: (2026)
by: Zhang, Hanxin, et al.
Published: (2026)
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
by: Chen, William, et al.
Published: (2026)
by: Chen, William, et al.
Published: (2026)
See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
by: Chen, Guangyan, et al.
Published: (2025)
by: Chen, Guangyan, et al.
Published: (2025)
Beyond Waypoints: Dual-Heatmap Grounding for Cross-Embodiment Semantic Navigation
by: Yun, Kaijie, et al.
Published: (2026)
by: Yun, Kaijie, et al.
Published: (2026)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
Similar Items
-
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
by: Xu, Kechun, et al.
Published: (2025) -
ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
by: Lou, Zhichen, et al.
Published: (2025) -
A Joint Modeling of Vision-Language-Action for Target-oriented Grasping in Clutter
by: Xu, Kechun, et al.
Published: (2023) -
Direction Matters: Learning Force Direction Enables Sim-to-Real Contact-Rich Manipulation
by: Yang, Yifei, et al.
Published: (2026) -
Grasp, See, and Place: Efficient Unknown Object Rearrangement with Policy Structure Prior
by: Xu, Kechun, et al.
Published: (2024)