VLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Beichen, Zhang, Juexiao, Dong, Shuwen, Fang, Irving, Feng, Chen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
di: Fang, Irving, et al.
Pubblicazione: (2025)
di: Fang, Irving, et al.
Pubblicazione: (2025)
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
di: Dai, Tingjun, et al.
Pubblicazione: (2026)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
di: Shao, Rui, et al.
Pubblicazione: (2025)
di: Shao, Rui, et al.
Pubblicazione: (2025)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025)
di: Li, Qixiu, et al.
Pubblicazione: (2025)
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)
What Matters in Building Vision-Language-Action Models for Generalist Robots
di: Li, Xinghang, et al.
Pubblicazione: (2024)
di: Li, Xinghang, et al.
Pubblicazione: (2024)
QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
di: Ding, Pengxiang, et al.
Pubblicazione: (2023)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
di: Zhou, Weijie, et al.
Pubblicazione: (2025)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
di: Wang, Hongyu, et al.
Pubblicazione: (2025)
DynamicVLA: A Vision-Language-Action Model for Dynamic Object Manipulation
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
di: Xie, Haozhe, et al.
Pubblicazione: (2026)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
di: Han, Xiaofeng, et al.
Pubblicazione: (2025)
di: Han, Xiaofeng, et al.
Pubblicazione: (2025)
ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
di: Song, Wenxuan, et al.
Pubblicazione: (2025)
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
di: Cheng, An-Chieh, et al.
Pubblicazione: (2024)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
di: Din, Muhayy Ud, et al.
Pubblicazione: (2025)
di: Din, Muhayy Ud, et al.
Pubblicazione: (2025)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
di: Kerr, Justin, et al.
Pubblicazione: (2024)
di: Kerr, Justin, et al.
Pubblicazione: (2024)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
di: Wen, Junjie, et al.
Pubblicazione: (2024)
di: Wen, Junjie, et al.
Pubblicazione: (2024)
See What Matters: Differentiable Grid Sample Pruning for Generalizable Vision-Language-Action Model
di: Feng, Yixu, et al.
Pubblicazione: (2026)
di: Feng, Yixu, et al.
Pubblicazione: (2026)
Seeing Realism from Simulation: Efficient Video Transfer for Vision-Language-Action Data Augmentation
di: Hui, Chenyu, et al.
Pubblicazione: (2026)
di: Hui, Chenyu, et al.
Pubblicazione: (2026)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
di: Liu, Mengzhen, et al.
Pubblicazione: (2026)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
di: Shi, Hao, et al.
Pubblicazione: (2025)
di: Shi, Hao, et al.
Pubblicazione: (2025)
MAP-VLA: Memory-Augmented Prompting for Vision-Language-Action Model in Robotic Manipulation
di: Li, Runhao, et al.
Pubblicazione: (2025)
di: Li, Runhao, et al.
Pubblicazione: (2025)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
di: Chopra, Samarth, et al.
Pubblicazione: (2025)
di: Chopra, Samarth, et al.
Pubblicazione: (2025)
Benchmarking Vision, Language, & Action Models on Robotic Learning Tasks
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
di: Guruprasad, Pranav, et al.
Pubblicazione: (2024)
This&That: Language-Gesture Controlled Video Generation for Robot Planning
di: Wang, Boyang, et al.
Pubblicazione: (2024)
di: Wang, Boyang, et al.
Pubblicazione: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
di: Zhang, Chubin, et al.
Pubblicazione: (2026)
di: Zhang, Chubin, et al.
Pubblicazione: (2026)
See and Switch: Vision-Based Branching for Interactive Robot-Skill Programming
di: Vanc, Petr, et al.
Pubblicazione: (2026)
di: Vanc, Petr, et al.
Pubblicazione: (2026)
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models
di: Bai, Yu, et al.
Pubblicazione: (2026)
di: Bai, Yu, et al.
Pubblicazione: (2026)
HRIBench: Benchmarking Vision-Language Models for Real-Time Human Perception in Human-Robot Interaction
di: Shi, Zhonghao, et al.
Pubblicazione: (2025)
di: Shi, Zhonghao, et al.
Pubblicazione: (2025)
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
di: Liu, Jiaxing, et al.
Pubblicazione: (2026)
di: Liu, Jiaxing, et al.
Pubblicazione: (2026)
MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
di: Huang, Ting, et al.
Pubblicazione: (2025)
di: Huang, Ting, et al.
Pubblicazione: (2025)
Observe Then Act: Asynchronous Active Vision-Action Model for Robotic Manipulation
di: Wang, Guokang, et al.
Pubblicazione: (2024)
di: Wang, Guokang, et al.
Pubblicazione: (2024)
Explainable Adversarial-Robust Vision-Language-Action Model for Robotic Manipulation
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
di: Kim, Ju-Young, et al.
Pubblicazione: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
di: Liu, Shuo, et al.
Pubblicazione: (2026)
di: Liu, Shuo, et al.
Pubblicazione: (2026)
ChatVLA: Unified Multimodal Understanding and Robot Control with Vision-Language-Action Model
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
di: Zhou, Zhongyi, et al.
Pubblicazione: (2025)
How Robot Dogs See the Unseeable: Improving Visual Interpretability via Peering for Exploratory Robots
di: Bimber, Oliver, et al.
Pubblicazione: (2025)
di: Bimber, Oliver, et al.
Pubblicazione: (2025)
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
di: Chen, Shizhe, et al.
Pubblicazione: (2025)
SIMPACT: Simulation-Enabled Action Planning using Vision-Language Models
di: Liu, Haowen, et al.
Pubblicazione: (2025)
di: Liu, Haowen, et al.
Pubblicazione: (2025)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
di: Xu, Kechun, et al.
Pubblicazione: (2025)
di: Xu, Kechun, et al.
Pubblicazione: (2025)
Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation
di: Ding, Hongyu, et al.
Pubblicazione: (2026)
di: Ding, Hongyu, et al.
Pubblicazione: (2026)
Documenti analoghi
-
From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
di: Fang, Irving, et al.
Pubblicazione: (2025) -
See, Plan, Rewind: Progress-Aware Vision-Language-Action Models for Robust Robotic Manipulation
di: Dai, Tingjun, et al.
Pubblicazione: (2026) -
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
di: Shao, Rui, et al.
Pubblicazione: (2025) -
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
di: Li, Qixiu, et al.
Pubblicazione: (2025) -
NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation
di: Zhang, Jiazhao, et al.
Pubblicazione: (2024)