Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
Fuente:
arXiv
Guardado en:
| Autores principales: | Garcia, Ricardo, Chen, Shizhe, Schmid, Cordelia |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
por: Chen, Shizhe, et al.
Publicado: (2025)
por: Chen, Shizhe, et al.
Publicado: (2025)
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
por: Pacaud, Paul, et al.
Publicado: (2025)
por: Pacaud, Paul, et al.
Publicado: (2025)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
por: Chen, Shizhe, et al.
Publicado: (2026)
por: Chen, Shizhe, et al.
Publicado: (2026)
Online 3D Scene Reconstruction Using Neural Object Priors
por: Chabal, Thomas, et al.
Publicado: (2025)
por: Chabal, Thomas, et al.
Publicado: (2025)
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
por: Chen, Zerui, et al.
Publicado: (2024)
por: Chen, Zerui, et al.
Publicado: (2024)
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
por: Chen, Zerui, et al.
Publicado: (2026)
por: Chen, Zerui, et al.
Publicado: (2026)
SUGAR: Pre-training 3D Visual Representations for Robotics
por: Chen, Shizhe, et al.
Publicado: (2024)
por: Chen, Shizhe, et al.
Publicado: (2024)
FOM-Nav: Frontier-Object Maps for Object Goal Navigation
por: Chabal, Thomas, et al.
Publicado: (2025)
por: Chabal, Thomas, et al.
Publicado: (2025)
Towards Generalizable Robotic Manipulation in Dynamic Environments
por: Fang, Heng, et al.
Publicado: (2026)
por: Fang, Heng, et al.
Publicado: (2026)
ArtiBench and ArtiBrain: Benchmarking Generalizable Vision-Language Articulated Object Manipulation
por: Wu, Yuhan, et al.
Publicado: (2025)
por: Wu, Yuhan, et al.
Publicado: (2025)
Generalizable Humanoid Manipulation with 3D Diffusion Policies
por: Ze, Yanjie, et al.
Publicado: (2024)
por: Ze, Yanjie, et al.
Publicado: (2024)
Learning Generalizable 3D Manipulation With 10 Demonstrations
por: Ren, Yu, et al.
Publicado: (2024)
por: Ren, Yu, et al.
Publicado: (2024)
MetricNet: Recovering Metric Scale in Generative Navigation Policies
por: Nayak, Abhijeet, et al.
Publicado: (2025)
por: Nayak, Abhijeet, et al.
Publicado: (2025)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
por: Khan, Zeeshan, et al.
Publicado: (2025)
por: Khan, Zeeshan, et al.
Publicado: (2025)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
por: Yin, Zhenhan, et al.
Publicado: (2025)
por: Yin, Zhenhan, et al.
Publicado: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
por: Liu, Mengzhen, et al.
Publicado: (2026)
por: Liu, Mengzhen, et al.
Publicado: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
por: Huang, Haifeng, et al.
Publicado: (2025)
por: Huang, Haifeng, et al.
Publicado: (2025)
O$^3$Afford: One-Shot 3D Object-to-Object Affordance Grounding for Generalizable Robotic Manipulation
por: Tian, Tongxuan, et al.
Publicado: (2025)
por: Tian, Tongxuan, et al.
Publicado: (2025)
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
por: Chen, Hanzhi, et al.
Publicado: (2025)
por: Chen, Hanzhi, et al.
Publicado: (2025)
ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
por: Li, Zheng, et al.
Publicado: (2025)
por: Li, Zheng, et al.
Publicado: (2025)
Contrastive Imitation Learning for Language-guided Multi-Task Robotic Manipulation
por: Ma, Teli, et al.
Publicado: (2024)
por: Ma, Teli, et al.
Publicado: (2024)
ManiSoft: Towards Vision-Language Manipulation for Soft Continuum Robotics
por: Wei, Ziyu, et al.
Publicado: (2026)
por: Wei, Ziyu, et al.
Publicado: (2026)
Fast Visuomotor Policy for Robotic Manipulation
por: Jia, Jingkai, et al.
Publicado: (2025)
por: Jia, Jingkai, et al.
Publicado: (2025)
TinyVLA: Towards Fast, Data-Efficient Vision-Language-Action Models for Robotic Manipulation
por: Wen, Junjie, et al.
Publicado: (2024)
por: Wen, Junjie, et al.
Publicado: (2024)
BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
por: Wang, Hongyu, et al.
Publicado: (2025)
por: Wang, Hongyu, et al.
Publicado: (2025)
Vision Language Action Models in Robotic Manipulation: A Systematic Review
por: Din, Muhayy Ud, et al.
Publicado: (2025)
por: Din, Muhayy Ud, et al.
Publicado: (2025)
Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
por: Shao, Rui, et al.
Publicado: (2025)
por: Shao, Rui, et al.
Publicado: (2025)
Robotic Programmer: Video Instructed Policy Code Generation for Robotic Manipulation
por: Xie, Senwei, et al.
Publicado: (2025)
por: Xie, Senwei, et al.
Publicado: (2025)
Generalized Robot 3D Vision-Language Model with Fast Rendering and Pre-Training Vision-Language Alignment
por: Liu, Kangcheng, et al.
Publicado: (2023)
por: Liu, Kangcheng, et al.
Publicado: (2023)
3D Dynamics-Aware Manipulation: Endowing Manipulation Policies with 3D Foresight
por: He, Yuxin, et al.
Publicado: (2025)
por: He, Yuxin, et al.
Publicado: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
por: Kuang, Yuxuan, et al.
Publicado: (2024)
por: Kuang, Yuxuan, et al.
Publicado: (2024)
LARY: A Latent Action Representation Yielding Benchmark for Generalizable Vision-to-Action Alignment
por: Nie, Dujun, et al.
Publicado: (2026)
por: Nie, Dujun, et al.
Publicado: (2026)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
por: Bousselham, Walid, et al.
Publicado: (2025)
por: Bousselham, Walid, et al.
Publicado: (2025)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
por: Ni, Minheng, et al.
Publicado: (2024)
por: Ni, Minheng, et al.
Publicado: (2024)
EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
por: Chopra, Samarth, et al.
Publicado: (2025)
por: Chopra, Samarth, et al.
Publicado: (2025)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
por: Lin, Haitao, et al.
Publicado: (2026)
por: Lin, Haitao, et al.
Publicado: (2026)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
por: Duan, Jiafei, et al.
Publicado: (2024)
por: Duan, Jiafei, et al.
Publicado: (2024)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation
por: Zhao, Hongxiang, et al.
Publicado: (2025)
por: Zhao, Hongxiang, et al.
Publicado: (2025)
3D-CDRGP: Towards Cross-Device Robotic Grasping Policy in 3D Open World
por: Zhao, Weiguang, et al.
Publicado: (2024)
por: Zhao, Weiguang, et al.
Publicado: (2024)
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
por: Qi, Yu, et al.
Publicado: (2025)
por: Qi, Yu, et al.
Publicado: (2025)
Ejemplares similares
-
Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
por: Chen, Shizhe, et al.
Publicado: (2025) -
Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation
por: Pacaud, Paul, et al.
Publicado: (2025) -
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
por: Chen, Shizhe, et al.
Publicado: (2026) -
Online 3D Scene Reconstruction Using Neural Object Priors
por: Chabal, Thomas, et al.
Publicado: (2025) -
ViViDex: Learning Vision-based Dexterous Manipulation from Human Videos
por: Chen, Zerui, et al.
Publicado: (2024)