RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Enshen, An, Jingkun, Chi, Cheng, Han, Yi, Rong, Shanyu, Zhang, Chi, Wang, Pengwei, Wang, Zhongyuan, Huang, Tiejun, Sheng, Lu, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025)
by: Zhou, Enshen, et al.
Published: (2025)
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025)
by: Han, Yi, et al.
Published: (2025)
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026)
by: Liu, Mengzhen, et al.
Published: (2026)
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
by: Tan, Huajie, et al.
Published: (2026)
by: Tan, Huajie, et al.
Published: (2026)
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
by: Zhou, Enshen, et al.
Published: (2024)
by: Zhou, Enshen, et al.
Published: (2024)
Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
RoboBrain 2.5: Depth in Sight, Time in Mind
by: Tan, Huajie, et al.
Published: (2026)
by: Tan, Huajie, et al.
Published: (2026)
Latent Reasoning VLA: Latent Thinking and Prediction for Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
Reshaping Action Error Distributions for Reliable Vision-Language-Action Models
by: Bai, Shuanghao, et al.
Published: (2026)
by: Bai, Shuanghao, et al.
Published: (2026)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
by: Tang, Yingbo, et al.
Published: (2025)
by: Tang, Yingbo, et al.
Published: (2025)
RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
by: Tan, Huajie, et al.
Published: (2025)
by: Tan, Huajie, et al.
Published: (2025)
RoboBrain 2.0 Technical Report
by: BAAI RoboBrain Team, et al.
Published: (2025)
by: BAAI RoboBrain Team, et al.
Published: (2025)
PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
by: Peng, Cheng, et al.
Published: (2025)
by: Peng, Cheng, et al.
Published: (2025)
RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete
by: Ji, Yuheng, et al.
Published: (2025)
by: Ji, Yuheng, et al.
Published: (2025)
MapNav: A Novel Memory Representation via Annotated Semantic Maps for Vision-and-Language Navigation
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
PRM-as-a-Judge: A Dense Evaluation Paradigm for Fine-Grained Robotic Auditing
by: Ji, Yuheng, et al.
Published: (2026)
by: Ji, Yuheng, et al.
Published: (2026)
VEGA: Visual Encoder Grounding Alignment for Spatially-Aware Vision-Language-Action Models
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
by: Fu, Yankai, et al.
Published: (2025)
by: Fu, Yankai, et al.
Published: (2025)
From Language to Locomotion: Retargeting-free Humanoid Control via Motion Latent Guidance
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
by: Li, Xiaoqi, et al.
Published: (2025)
by: Li, Xiaoqi, et al.
Published: (2025)
RoboArmGS: High-Quality Robotic Arm Splatting via Bézier Curve Refinement
by: Wang, Hao, et al.
Published: (2025)
by: Wang, Hao, et al.
Published: (2025)
Do You Have Freestyle? Expressive Humanoid Locomotion via Audio Control
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
RoboCOIN: An Open-Sourced Bimanual Robotic Data Collection for Integrated Manipulation
by: Wu, Shihan, et al.
Published: (2025)
by: Wu, Shihan, et al.
Published: (2025)
Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots
by: Zou, Ziqing, et al.
Published: (2026)
by: Zou, Ziqing, et al.
Published: (2026)
$NavA^3$: Understanding Any Instruction, Navigating Anywhere, Finding Anything
by: Zhang, Lingfeng, et al.
Published: (2025)
by: Zhang, Lingfeng, et al.
Published: (2025)
Mask World Model: Predicting What Matters for Robust Robot Policy Learning
by: Lou, Yunfan, et al.
Published: (2026)
by: Lou, Yunfan, et al.
Published: (2026)
RoboStream: Weaving Spatio-Temporal Reasoning with Memory in Vision-Language Models for Robotics
by: Huang, Yuzhi, et al.
Published: (2026)
by: Huang, Yuzhi, et al.
Published: (2026)
RoboFlamingo-Plus: Fusion of Depth and RGB Perception with Vision-Language Models for Enhanced Robotic Manipulation
by: Wang, Sheng
Published: (2025)
by: Wang, Sheng
Published: (2025)
RoboPaint: From Human Demonstration to Any Robot and Any View
by: Fan, Jiacheng, et al.
Published: (2026)
by: Fan, Jiacheng, et al.
Published: (2026)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
Whole-body Humanoid Robot Locomotion with Human Reference
by: Zhang, Qiang, et al.
Published: (2024)
by: Zhang, Qiang, et al.
Published: (2024)
RoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot Learning
by: Geng, Haoran, et al.
Published: (2025)
by: Geng, Haoran, et al.
Published: (2025)
OmniUMI: Towards Physically Grounded Robot Learning via Human-Aligned Multimodal Interaction
by: Luo, Shaqi, et al.
Published: (2026)
by: Luo, Shaqi, et al.
Published: (2026)
RoboPocket: Improve Robot Policies Instantly with Your Phone
by: Fang, Junjie, et al.
Published: (2026)
by: Fang, Junjie, et al.
Published: (2026)
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
by: Wang, Sicheng, et al.
Published: (2025)
by: Wang, Sicheng, et al.
Published: (2025)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
by: Ji, Yuheng, et al.
Published: (2025)
by: Ji, Yuheng, et al.
Published: (2025)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
by: Zhang, Rongyu, et al.
Published: (2025)
by: Zhang, Rongyu, et al.
Published: (2025)
RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark
by: Lei, Huashuo, et al.
Published: (2026)
by: Lei, Huashuo, et al.
Published: (2026)
Similar Items
-
RoboTracer: Mastering Spatial Trace with Reasoning in Vision-Language Models for Robotics
by: Zhou, Enshen, et al.
Published: (2025) -
TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
by: Han, Yi, et al.
Published: (2025) -
SaPaVe: Towards Active Perception and Manipulation in Vision-Language-Action Models for Robotics
by: Liu, Mengzhen, et al.
Published: (2026) -
Action-Sketcher: From Reasoning to Action via Visual Sketches for Long-Horizon Robotic Manipulation
by: Tan, Huajie, et al.
Published: (2026) -
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
by: Zhou, Enshen, et al.
Published: (2024)