Survey of Vision-Language-Action Models for Embodied Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Haoran, Chen, Yuhui, Cui, Wenbo, Liu, Weiheng, Liu, Kai, Zhou, Mingcai, Zhang, Zhengtao, Zhao, Dongbin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
von: Li, Yixuan, et al.
Veröffentlicht: (2025)
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
von: Jiang, Zhennan, et al.
Veröffentlicht: (2025)
von: Jiang, Zhennan, et al.
Veröffentlicht: (2025)
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
von: Cui, Wenbo, et al.
Veröffentlicht: (2025)
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
von: Chen, Yuhui, et al.
Veröffentlicht: (2025)
von: Chen, Yuhui, et al.
Veröffentlicht: (2025)
TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning
von: Chen, Yuhui, et al.
Veröffentlicht: (2025)
von: Chen, Yuhui, et al.
Veröffentlicht: (2025)
DiffuDepGrasp: Diffusion-based Depth Noise Modeling Empowers Sim2Real Robotic Grasping
von: Zhou, Yingting, et al.
Veröffentlicht: (2025)
von: Zhou, Yingting, et al.
Veröffentlicht: (2025)
GAPartManip: A Large-scale Part-centric Dataset for Material-Agnostic Articulated Object Manipulation
von: Cui, Wenbo, et al.
Veröffentlicht: (2024)
von: Cui, Wenbo, et al.
Veröffentlicht: (2024)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
VGGT-DP: Generalizable Robot Control via Vision Foundation Models
von: Ge, Shijia, et al.
Veröffentlicht: (2025)
von: Ge, Shijia, et al.
Veröffentlicht: (2025)
A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
von: Zhai, Shaopeng, et al.
Veröffentlicht: (2025)
Experiences from Benchmarking Vision-Language-Action Models for Robotic Manipulation
von: Zhang, Yihao, et al.
Veröffentlicht: (2025)
von: Zhang, Yihao, et al.
Veröffentlicht: (2025)
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
von: Pei, Xiaohuan, et al.
Veröffentlicht: (2025)
Towards Long-horizon Embodied Agents with Tool-Aligned Vision-Language-Action Models
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
von: Lei, Zixing, et al.
Veröffentlicht: (2026)
Grounding Sim-to-Real Generalization in Dexterous Manipulation: An Empirical Study with Vision-Language-Action Models
von: Jin, Ruixing, et al.
Veröffentlicht: (2026)
von: Jin, Ruixing, et al.
Veröffentlicht: (2026)
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
von: Zou, Teqiang, et al.
Veröffentlicht: (2025)
von: Zou, Teqiang, et al.
Veröffentlicht: (2025)
LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
von: Miao, Cui, et al.
Veröffentlicht: (2025)
von: Miao, Cui, et al.
Veröffentlicht: (2025)
OccVLA: Vision-Language-Action Model with Implicit 3D Occupancy Supervision
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
von: Liu, Ruixun, et al.
Veröffentlicht: (2025)
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
von: Hong, Youngjin, et al.
Veröffentlicht: (2025)
von: Hong, Youngjin, et al.
Veröffentlicht: (2025)
Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach
von: Yang, Siyuan, et al.
Veröffentlicht: (2025)
von: Yang, Siyuan, et al.
Veröffentlicht: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
von: Liu, Jinkun, et al.
Veröffentlicht: (2026)
DeMaVLA: A Vision-Language-Action Foundation Model for Generalizable Deformable Manipulation
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
von: Su, Taiyi, et al.
Veröffentlicht: (2026)
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
InCoM: Intent-Driven Perception and Structured Coordination for Mobile Manipulation
von: Liu, Jiahao, et al.
Veröffentlicht: (2026)
von: Liu, Jiahao, et al.
Veröffentlicht: (2026)
DexHiL: A Human-in-the-Loop Framework for Vision-Language-Action Model Post-Training in Dexterous Manipulation
von: Han, Yifan, et al.
Veröffentlicht: (2026)
von: Han, Yifan, et al.
Veröffentlicht: (2026)
Vision-Language-Action in Robotics: A Survey of Datasets, Benchmarks, and Data Engines
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
von: Wang, Ziyao, et al.
Veröffentlicht: (2026)
X-DiffVLA: X-Embodied Diffusion Action Heads for Vision-Language-Action Models
von: Li, Boyu, et al.
Veröffentlicht: (2026)
von: Li, Boyu, et al.
Veröffentlicht: (2026)
A Survey of Robotic Navigation and Manipulation with Physics Simulators in the Era of Embodied AI
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
von: Wong, Lik Hang Kenny, et al.
Veröffentlicht: (2025)
ALOE: Action-Level Off-Policy Evaluation for Vision-Language-Action Model Post-Training
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
von: Yang, Rushuai, et al.
Veröffentlicht: (2026)
WoVR: World Models as Reliable Simulators for Post-Training VLA Policies with RL
von: Jiang, Zhennan, et al.
Veröffentlicht: (2026)
von: Jiang, Zhennan, et al.
Veröffentlicht: (2026)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Chen, Xiaoyu, et al.
Veröffentlicht: (2025)
A Survey on Robotics with Foundation Models: toward Embodied AI
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Xu, Zhiyuan, et al.
Veröffentlicht: (2024)
RoboGPT: an intelligent agent of making embodied long-term decisions for daily instruction tasks
von: Chen, Yaran, et al.
Veröffentlicht: (2023)
von: Chen, Yaran, et al.
Veröffentlicht: (2023)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
von: Takagi, Yusuke, et al.
Veröffentlicht: (2026)
von: Takagi, Yusuke, et al.
Veröffentlicht: (2026)
Characterizing Vision-Language-Action Models across XPUs: Constraints and Acceleration for On-Robot Deployment
von: Zhou, Kaijun, et al.
Veröffentlicht: (2026)
von: Zhou, Kaijun, et al.
Veröffentlicht: (2026)
Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
von: Zhang, Yi, et al.
Veröffentlicht: (2026)
Octopus: Embodied Vision-Language Programmer from Environmental Feedback
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
von: Yang, Jingkang, et al.
Veröffentlicht: (2023)
Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
von: Yuan, Yifu, et al.
Veröffentlicht: (2025)
von: Yuan, Yifu, et al.
Veröffentlicht: (2025)
EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment
von: Gao, Chen, et al.
Veröffentlicht: (2024)
von: Gao, Chen, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
von: Li, Yixuan, et al.
Veröffentlicht: (2025) -
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
von: Jiang, Zhennan, et al.
Veröffentlicht: (2025) -
CL3R: 3D Reconstruction and Contrastive Learning for Enhanced Robotic Manipulation Representations
von: Cui, Wenbo, et al.
Veröffentlicht: (2025) -
ConRFT: A Reinforced Fine-tuning Method for VLA Models via Consistency Policy
von: Chen, Yuhui, et al.
Veröffentlicht: (2025) -
TeViR: Text-to-Video Reward with Diffusion Models for Efficient Reinforcement Learning
von: Chen, Yuhui, et al.
Veröffentlicht: (2025)