Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Song, Zijian, Li, Qichang, Zhou, Jiawei, Yuan, Zhenlong, Chen, Tianshui, Lin, Liang, Wang, Guangrun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
von: Song, Zijian, et al.
Veröffentlicht: (2026)
von: Song, Zijian, et al.
Veröffentlicht: (2026)
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
von: Vuong, An Dinh, et al.
Veröffentlicht: (2025)
von: Vuong, An Dinh, et al.
Veröffentlicht: (2025)
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
von: Song, Zijian, et al.
Veröffentlicht: (2025)
von: Song, Zijian, et al.
Veröffentlicht: (2025)
Look-to-Touch: A Vision-Enhanced Proximity and Tactile Sensor for Distance and Geometry Perception in Robotic Manipulation
von: Dong, Yueshi, et al.
Veröffentlicht: (2025)
von: Dong, Yueshi, et al.
Veröffentlicht: (2025)
Implicit Geometry Representations for Vision-and-Language Navigation from Web Videos
von: Han, Mingfei, et al.
Veröffentlicht: (2026)
von: Han, Mingfei, et al.
Veröffentlicht: (2026)
Stable Language Guidance for Vision-Language-Action Models
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
von: Zhan, Zhihao, et al.
Veröffentlicht: (2026)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
Geometry-aware 4D Video Generation for Robot Manipulation
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
VLMPC: Vision-Language Model Predictive Control for Robotic Manipulation
von: Zhao, Wentao, et al.
Veröffentlicht: (2024)
von: Zhao, Wentao, et al.
Veröffentlicht: (2024)
GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
von: Abouzeid, Ali, et al.
Veröffentlicht: (2025)
VLATest: Testing and Evaluating Vision-Language-Action Models for Robotic Manipulation
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
Task-oriented Robotic Manipulation with Vision Language Models
von: Guran, Nurhan Bulus, et al.
Veröffentlicht: (2024)
von: Guran, Nurhan Bulus, et al.
Veröffentlicht: (2024)
CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
von: An, Zijian, et al.
Veröffentlicht: (2025)
von: An, Zijian, et al.
Veröffentlicht: (2025)
Threading Optimization for Vision-Language-Action Model Inference in Low-Cost Smart Agricultural Manipulation
von: Truongcao, Keith, et al.
Veröffentlicht: (2026)
von: Truongcao, Keith, et al.
Veröffentlicht: (2026)
DyGeoVLN: Infusing Dynamic Geometry Foundation Model into Vision-Language Navigation
von: Liu, Xiangchen, et al.
Veröffentlicht: (2026)
von: Liu, Xiangchen, et al.
Veröffentlicht: (2026)
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation
von: Zuo, Kuangji, et al.
Veröffentlicht: (2026)
von: Zuo, Kuangji, et al.
Veröffentlicht: (2026)
LADEV: A Language-Driven Testing and Evaluation Platform for Vision-Language-Action Models in Robotic Manipulation
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
von: Wang, Zhijie, et al.
Veröffentlicht: (2024)
Online Robot Navigation and Manipulation with Distilled Vision-Language Models
von: Liu, Kangcheng
Veröffentlicht: (2024)
von: Liu, Kangcheng
Veröffentlicht: (2024)
TAG: Target-Agnostic Guidance for Stable Object-Centric Inference in Vision-Language-Action Models
von: Zhou, Jiaying, et al.
Veröffentlicht: (2026)
von: Zhou, Jiaying, et al.
Veröffentlicht: (2026)
Pure Vision Language Action (VLA) Models: A Comprehensive Survey
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
von: Zhang, Dapeng, et al.
Veröffentlicht: (2025)
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
von: Liu, Minghuan, et al.
Veröffentlicht: (2025)
von: Liu, Minghuan, et al.
Veröffentlicht: (2025)
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
von: Tang, Yikai, et al.
Veröffentlicht: (2025)
von: Tang, Yikai, et al.
Veröffentlicht: (2025)
MSGField: A Unified Scene Representation Integrating Motion, Semantics, and Geometry for Robotic Manipulation
von: Sheng, Yu, et al.
Veröffentlicht: (2024)
von: Sheng, Yu, et al.
Veröffentlicht: (2024)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
von: Li, Qixiu, et al.
Veröffentlicht: (2025)
Representing Robot Geometry as Distance Fields: Applications to Whole-body Manipulation
von: Li, Yiming, et al.
Veröffentlicht: (2023)
von: Li, Yiming, et al.
Veröffentlicht: (2023)
AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
Confusion-Aware In-Context-Learning for Vision-Language Models in Robotic Manipulation
von: He, Yayun, et al.
Veröffentlicht: (2026)
von: He, Yayun, et al.
Veröffentlicht: (2026)
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
von: Zhou, Huayi, et al.
Veröffentlicht: (2025)
von: Zhou, Huayi, et al.
Veröffentlicht: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
von: Vo, Khoa, et al.
Veröffentlicht: (2025)
von: Vo, Khoa, et al.
Veröffentlicht: (2025)
CrayonRobo: Object-Centric Prompt-Driven Vision-Language-Action Model for Robotic Manipulation
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
von: Zhao, Wei, et al.
Veröffentlicht: (2025)
MARVL: Multi-Stage Guidance for Robotic Manipulation via Vision-Language Models
von: Zhou, Xunlan, et al.
Veröffentlicht: (2026)
von: Zhou, Xunlan, et al.
Veröffentlicht: (2026)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
Vision-Guided Loco-Manipulation with a Snake Robot
von: Salagame, Adarsh, et al.
Veröffentlicht: (2025)
von: Salagame, Adarsh, et al.
Veröffentlicht: (2025)
VIP: Vision Instructed Pre-training for Robotic Manipulation
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
von: Li, Zhuoling, et al.
Veröffentlicht: (2024)
MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
von: Zhang, Rongyu, et al.
Veröffentlicht: (2025)
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
von: Shi, Hao, et al.
Veröffentlicht: (2025)
von: Shi, Hao, et al.
Veröffentlicht: (2025)
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
von: Duan, Jiafei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
von: Song, Zijian, et al.
Veröffentlicht: (2026) -
Improving Robotic Manipulation with Efficient Geometry-Aware Vision Encoder
von: Vuong, An Dinh, et al.
Veröffentlicht: (2025) -
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
von: Chen, Yuhao, et al.
Veröffentlicht: (2026) -
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
von: Song, Zijian, et al.
Veröffentlicht: (2025) -
Look-to-Touch: A Vision-Enhanced Proximity and Tactile Sensor for Distance and Geometry Perception in Robotic Manipulation
von: Dong, Yueshi, et al.
Veröffentlicht: (2025)