Point What You Mean: Visually Grounded Instruction Policy
Fuente:
arXiv
Saved in:
| Main Authors: | Yu, Hang, Zhao, Juntu, Liu, Yufeng, Li, Kaiyu, Ma, Cheng, Zhang, Di, Hu, Yingdong, Chen, Guang, Xie, Junyuan, Guo, Junliang, Zhao, Junqiao, Gao, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Do You Need Proprioceptive States in Visuomotor Policies?
by: Zhao, Juntu, et al.
Published: (2025)
by: Zhao, Juntu, et al.
Published: (2025)
Learning Native Continuation for Action Chunking Flow Policies
by: Liu, Yufeng, et al.
Published: (2026)
by: Liu, Yufeng, et al.
Published: (2026)
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
by: Zhang, Di, et al.
Published: (2026)
by: Zhang, Di, et al.
Published: (2026)
KineDex: Learning Tactile-Informed Visuomotor Policies via Kinesthetic Teaching for Dexterous Manipulation
by: Zhang, Di, et al.
Published: (2025)
by: Zhang, Di, et al.
Published: (2025)
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Focus On What Matters: Separated Models For Visual-Based RL Generalization
by: Zhang, Di, et al.
Published: (2024)
by: Zhang, Di, et al.
Published: (2024)
Conditioning Matters: Training Diffusion Policies is Faster Than You Think
by: Dong, Zibin, et al.
Published: (2025)
by: Dong, Zibin, et al.
Published: (2025)
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
by: Lin, Fanqi, et al.
Published: (2025)
by: Lin, Fanqi, et al.
Published: (2025)
Grasp What You Want: Embodied Dexterous Grasping System Driven by Your Voice
by: Li, Junliang, et al.
Published: (2024)
by: Li, Junliang, et al.
Published: (2024)
Efficient Multi-Robot Motion Planning for Manifold-Constrained Manipulators by Randomized Scheduling and Informed Path Generation
by: Guo, Weihang, et al.
Published: (2024)
by: Guo, Weihang, et al.
Published: (2024)
UNO Push: Unified Nonprehensile Object Pushing via Non-Parametric Estimation and Model Predictive Control
by: Wang, Gaotian, et al.
Published: (2024)
by: Wang, Gaotian, et al.
Published: (2024)
DL-SLOT: Dynamic LiDAR SLAM and object tracking based on collaborative graph optimization
by: Tian, Xuebo, et al.
Published: (2022)
by: Tian, Xuebo, et al.
Published: (2022)
Self-Imitated Diffusion Policy for Efficient and Robust Visual Navigation
by: Zhang, Runhua, et al.
Published: (2026)
by: Zhang, Runhua, et al.
Published: (2026)
CoPa: General Robotic Manipulation through Spatial Constraints of Parts with Foundation Models
by: Huang, Haoxu, et al.
Published: (2024)
by: Huang, Haoxu, et al.
Published: (2024)
MotionTrans: Human VR Data Enable Motion-Level Learning for Robotic Manipulation Policies
by: Yuan, Chengbo, et al.
Published: (2025)
by: Yuan, Chengbo, et al.
Published: (2025)
Convex Hull-based Algebraic Constraint for Visual Quadric SLAM
by: Yu, Xiaolong, et al.
Published: (2025)
by: Yu, Xiaolong, et al.
Published: (2025)
Wearable Roller Rings to Augment In-Hand Manipulation through Active Surfaces
by: Webb, Hayden, et al.
Published: (2024)
by: Webb, Hayden, et al.
Published: (2024)
Caging in Time: A Framework for Robust Object Manipulation under Uncertainties and Limited Robot Perception
by: Wang, Gaotian, et al.
Published: (2024)
by: Wang, Gaotian, et al.
Published: (2024)
Collision-Inclusive Manipulation Planning for Occluded Object Grasping via Compliant Robot Motions
by: Ren, Kejia, et al.
Published: (2024)
by: Ren, Kejia, et al.
Published: (2024)
ManiDreams: An Open-Source Library for Robust Object Manipulation via Uncertainty-aware Task-specific Intuitive Physics
by: Wang, Gaotian, et al.
Published: (2026)
by: Wang, Gaotian, et al.
Published: (2026)
Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching
by: Ren, Kejia, et al.
Published: (2026)
by: Ren, Kejia, et al.
Published: (2026)
ShanghaiTech Mapping Robot is All You Need: Robot System for Collecting Universal Ground Vehicle Datasets
by: Xu, Bowen, et al.
Published: (2024)
by: Xu, Bowen, et al.
Published: (2024)
See What I Mean? Expressiveness and Clarity in Robot Display Design
by: Ebisu, Matthew, et al.
Published: (2025)
by: Ebisu, Matthew, et al.
Published: (2025)
GroundSLAM: A Robust Visual SLAM System for Warehouse Robots Using Ground Textures
by: Xu, Kuan, et al.
Published: (2017)
by: Xu, Kuan, et al.
Published: (2017)
N$^{3}$-Mapping: Normal Guided Neural Non-Projective Signed Distance Fields for Large-scale 3D Mapping
by: Song, Shuangfu, et al.
Published: (2024)
by: Song, Shuangfu, et al.
Published: (2024)
LIMOT: A Tightly-Coupled System for LiDAR-Inertial Odometry and Multi-Object Tracking
by: Zhu, Zhongyang, et al.
Published: (2023)
by: Zhu, Zhongyang, et al.
Published: (2023)
Contact-Grounded Policy: Dexterous Visuotactile Policy with Generative Contact Grounding
by: Xu, Zhengtong, et al.
Published: (2026)
by: Xu, Zhengtong, et al.
Published: (2026)
B4P: Simultaneous Grasp and Motion Planning for Object Placement via Parallelized Bidirectional Forests and Path Repair
by: Leebron, Benjamin H., et al.
Published: (2025)
by: Leebron, Benjamin H., et al.
Published: (2025)
VAMPO: Policy Optimization for Improving Visual Dynamics in Video Action Models
by: Ge, Zirui, et al.
Published: (2026)
by: Ge, Zirui, et al.
Published: (2026)
Data Scaling Laws in Imitation Learning for Robotic Manipulation
by: Hu, Yingdong, et al.
Published: (2024)
by: Hu, Yingdong, et al.
Published: (2024)
AirCrab: A Hybrid Aerial-Ground Manipulator with An Active Wheel
by: Cao, Muqing, et al.
Published: (2024)
by: Cao, Muqing, et al.
Published: (2024)
SAMP: Spatial Anchor-based Motion Policy for Collision-Aware Robotic Manipulators
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
A Visual Reinforcement Learning-Based Separate Primitive Policy for Peg-in-Hole Tasks
by: Xu, Zichun, et al.
Published: (2025)
by: Xu, Zichun, et al.
Published: (2025)
Augmented Reality for RObots (ARRO): Pointing Visuomotor Policies Towards Visual Robustness
by: Mirjalili, Reihaneh, et al.
Published: (2025)
by: Mirjalili, Reihaneh, et al.
Published: (2025)
What to Ignore, What to React: Visually Robust RL Fine-Tuning of VLA Models
by: Peng, Yuanfang, et al.
Published: (2026)
by: Peng, Yuanfang, et al.
Published: (2026)
LOG-LIO2: A LiDAR-Inertial Odometry with Efficient Uncertainty Analysis
by: Huang, Kai, et al.
Published: (2024)
by: Huang, Kai, et al.
Published: (2024)
VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning
by: Zhou, Tianxing, et al.
Published: (2026)
by: Zhou, Tianxing, et al.
Published: (2026)
Imitation Learning from Observation with Automatic Discount Scheduling
by: Liu, Yuyang, et al.
Published: (2023)
by: Liu, Yuyang, et al.
Published: (2023)
PL-VIWO2: A Lightweight, Fast and Robust Visual-Inertial-Wheel Odometry Using Points and Lines
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
PL-VIWO: A Lightweight and Robust Point-Line Monocular Visual Inertial Wheel Odometry
by: Zhang, Zhixin, et al.
Published: (2025)
by: Zhang, Zhixin, et al.
Published: (2025)
Similar Items
-
Do You Need Proprioceptive States in Visuomotor Policies?
by: Zhao, Juntu, et al.
Published: (2025) -
Learning Native Continuation for Action Chunking Flow Policies
by: Liu, Yufeng, et al.
Published: (2026) -
Learning Geometrically-Grounded 3D Visual Representations for View-Generalizable Robotic Manipulation
by: Zhang, Di, et al.
Published: (2026) -
KineDex: Learning Tactile-Informed Visuomotor Policies via Kinesthetic Teaching for Dexterous Manipulation
by: Zhang, Di, et al.
Published: (2025) -
ACSAC: Adaptive Chunk Size Actor-Critic with Causal Transformer Q-Network
by: Chen, Qian, et al.
Published: (2026)