RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Yuan, Wentao, Duan, Jiafei, Blukis, Valts, Pumacay, Wilbert, Krishna, Ranjay, Murali, Adithyavairavan, Mousavian, Arsalan, Fox, Dieter |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024)
by: Pumacay, Wilbert, et al.
Published: (2024)
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025)
by: Fang, Haoquan, et al.
Published: (2025)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024)
by: Duan, Jiafei, et al.
Published: (2024)
DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning
by: Huang, Huang, et al.
Published: (2024)
by: Huang, Huang, et al.
Published: (2024)
RoboEval: Where Robotic Manipulation Meets Structured and Scalable Evaluation
by: Wang, Yi Ru, et al.
Published: (2025)
by: Wang, Yi Ru, et al.
Published: (2025)
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
by: Song, Chan Hee, et al.
Published: (2024)
by: Song, Chan Hee, et al.
Published: (2024)
cuRoboV2: Dynamics-Aware Motion Generation with Depth-Fused Distance Fields for High-DoF Robots
by: Sundaralingam, Balakumar, et al.
Published: (2026)
by: Sundaralingam, Balakumar, et al.
Published: (2026)
EVE: Enabling Anyone to Train Robots using Augmented Reality
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
by: Lin, Zijun, et al.
Published: (2025)
by: Lin, Zijun, et al.
Published: (2025)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
OG-VLA: Orthographic Image Generation for 3D-Aware Vision-Language Action Model
by: Singh, Ishika, et al.
Published: (2025)
by: Singh, Ishika, et al.
Published: (2025)
Recurrent-Depth VLA: Implicit Test-Time Compute Scaling of Vision-Language-Action Models via Latent Iterative Reasoning
by: Tur, Yalcin, et al.
Published: (2026)
by: Tur, Yalcin, et al.
Published: (2026)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
RoboMD: Uncovering Robot Vulnerabilities through Semantic Potential Fields
by: Sagar, Som, et al.
Published: (2024)
by: Sagar, Som, et al.
Published: (2024)
PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
by: Huang, Wenlong, et al.
Published: (2026)
by: Huang, Wenlong, et al.
Published: (2026)
RVT-2: Learning Precise Manipulation from Few Demonstrations
by: Goyal, Ankit, et al.
Published: (2024)
by: Goyal, Ankit, et al.
Published: (2024)
RoboCade: Gamifying Robot Data Collection
by: Mirchandani, Suvir, et al.
Published: (2025)
by: Mirchandani, Suvir, et al.
Published: (2025)
Constrained Generative Sampling of 6-DoF Grasps
by: Lundell, Jens, et al.
Published: (2023)
by: Lundell, Jens, et al.
Published: (2023)
Slot-Level Robotic Placement via Visual Imitation from Single Human Video
by: Shan, Dandan, et al.
Published: (2025)
by: Shan, Dandan, et al.
Published: (2025)
Neural Implicit Representation for Building Digital Twins of Unknown Articulated Objects
by: Weng, Yijia, et al.
Published: (2024)
by: Weng, Yijia, et al.
Published: (2024)
TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics
by: Chen, Shirui, et al.
Published: (2026)
by: Chen, Shirui, et al.
Published: (2026)
GRS: Generating Robotic Simulation Tasks from Real-World Images
by: Zook, Alex, et al.
Published: (2024)
by: Zook, Alex, et al.
Published: (2024)
VLA-0: Building State-of-the-Art VLAs with Zero Modification
by: Goyal, Ankit, et al.
Published: (2025)
by: Goyal, Ankit, et al.
Published: (2025)
MolmoAct: Action Reasoning Models that can Reason in Space
by: Lee, Jason, et al.
Published: (2025)
by: Lee, Jason, et al.
Published: (2025)
GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training
by: Murali, Adithyavairavan, et al.
Published: (2025)
by: Murali, Adithyavairavan, et al.
Published: (2025)
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
by: Hao, Xiaoshuai, et al.
Published: (2025)
by: Hao, Xiaoshuai, et al.
Published: (2025)
RoboBlockly Studio: Conversational Block Programming with Embodied Robot Feedback for Computational Thinking
by: Li, Leyi, et al.
Published: (2026)
by: Li, Leyi, et al.
Published: (2026)
BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
by: Bhat, Vineet, et al.
Published: (2025)
by: Bhat, Vineet, et al.
Published: (2025)
Grasp-MPC: Closed-Loop Visual Grasping via Value-Guided Model Predictive Control
by: Yamada, Jun, et al.
Published: (2025)
by: Yamada, Jun, et al.
Published: (2025)
MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
by: Deshpande, Abhay, et al.
Published: (2026)
by: Deshpande, Abhay, et al.
Published: (2026)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
by: Xiao, Zhanqi, et al.
Published: (2026)
by: Xiao, Zhanqi, et al.
Published: (2026)
More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
by: Shao, Xinyu, et al.
Published: (2025)
by: Shao, Xinyu, et al.
Published: (2025)
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
by: Kim, Yejin, et al.
Published: (2026)
by: Kim, Yejin, et al.
Published: (2026)
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboPlayground: Democratizing Robotic Evaluation through Structured Physical Domains
by: Wang, Yi Ru, et al.
Published: (2026)
by: Wang, Yi Ru, et al.
Published: (2026)
3D FlowMatch Actor: Unified 3D Policy for Single- and Dual-Arm Manipulation
by: Gkanatsios, Nikolaos, et al.
Published: (2025)
by: Gkanatsios, Nikolaos, et al.
Published: (2025)
RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
RoboHitch: Learning Visual Affordance from Disordered Keypoints for Hitch Knots Tying
by: Zuo, Jiahui, et al.
Published: (2026)
by: Zuo, Jiahui, et al.
Published: (2026)
Similar Items
-
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
by: Duan, Jiafei, et al.
Published: (2024) -
THE COLOSSEUM: A Benchmark for Evaluating Generalization for Robotic Manipulation
by: Pumacay, Wilbert, et al.
Published: (2024) -
SAM2Act: Integrating Visual Foundation Model with A Memory Architecture for Robotic Manipulation
by: Fang, Haoquan, et al.
Published: (2025) -
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
by: Duan, Jiafei, et al.
Published: (2024) -
DiffusionSeeder: Seeding Motion Optimization with Diffusion for Rapid Motion Planning
by: Huang, Huang, et al.
Published: (2024)