GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Qian, Jingjing, Han, Boyao, Shi, Chen, Xiao, Lei, Yang, Long, Shi, Shaoshuai, Jiang, Li
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866915921092673536
author Qian, Jingjing
Han, Boyao
Shi, Chen
Xiao, Lei
Yang, Long
Shi, Shaoshuai
Jiang, Li
author_facet Qian, Jingjing
Han, Boyao
Shi, Chen
Xiao, Lei
Yang, Long
Shi, Shaoshuai
Jiang, Li
contents Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware VLA framework that augments a continuous-action policy with predictive kinematic and geometric priors. GeoPredict introduces a trajectory-level module that encodes motion history and predicts multi-step 3D keypoint trajectories of robot arms, and a predictive 3D Gaussian geometry module that forecasts workspace geometry with track-guided refinement along future keypoint trajectories. These predictive modules serve exclusively as training-time supervision through depth-based rendering, while inference requires only lightweight additional query tokens without invoking any 3D decoding. Experiments on RoboCasa Human-50, LIBERO, and real-world manipulation tasks show that GeoPredict consistently outperforms strong VLA baselines, especially in geometry-intensive and spatially demanding scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2512_16811
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
Qian, Jingjing
Han, Boyao
Shi, Chen
Xiao, Lei
Yang, Long
Shi, Shaoshuai
Jiang, Li
Computer Vision and Pattern Recognition
Robotics
Vision-Language-Action (VLA) models achieve strong generalization in robotic manipulation but remain largely reactive and 2D-centric, making them unreliable in tasks that require precise 3D reasoning. We propose GeoPredict, a geometry-aware VLA framework that augments a continuous-action policy with predictive kinematic and geometric priors. GeoPredict introduces a trajectory-level module that encodes motion history and predicts multi-step 3D keypoint trajectories of robot arms, and a predictive 3D Gaussian geometry module that forecasts workspace geometry with track-guided refinement along future keypoint trajectories. These predictive modules serve exclusively as training-time supervision through depth-based rendering, while inference requires only lightweight additional query tokens without invoking any 3D decoding. Experiments on RoboCasa Human-50, LIBERO, and real-world manipulation tasks show that GeoPredict consistently outperforms strong VLA baselines, especially in geometry-intensive and spatially demanding scenarios.
title GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2512.16811