GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhou, Kaichen, Chen, Yuzhen, Zhan, Fangneng, Hua, Hang, Chen, Grace, Chang, Xinhai, Qu, Ao, Du, Yilun, Liu, Zhuang, Liang, Paul Pu, Wang, Mengyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
von: Zhou, Kaichen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2025)
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
von: Chang, Xinhai, et al.
Veröffentlicht: (2024)
von: Chang, Xinhai, et al.
Veröffentlicht: (2024)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
von: Gu, Renjie, et al.
Veröffentlicht: (2026)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
von: Lillemark, Hansen Jin, et al.
Veröffentlicht: (2026)
von: Lillemark, Hansen Jin, et al.
Veröffentlicht: (2026)
RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations
von: Zhou, Kaichen, et al.
Veröffentlicht: (2024)
von: Zhou, Kaichen, et al.
Veröffentlicht: (2024)
Memorization in 3D Shape Generation: An Empirical Study
von: Pu, Shu, et al.
Veröffentlicht: (2025)
von: Pu, Shu, et al.
Veröffentlicht: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
von: Zhao, Minda, et al.
Veröffentlicht: (2026)
von: Zhao, Minda, et al.
Veröffentlicht: (2026)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
von: Song, Zijian, et al.
Veröffentlicht: (2026)
von: Song, Zijian, et al.
Veröffentlicht: (2026)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
von: Zhou, Pengfei, et al.
Veröffentlicht: (2026)
von: Zhou, Pengfei, et al.
Veröffentlicht: (2026)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanhan, et al.
Veröffentlicht: (2024)
SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds
von: Hou, Zhixing, et al.
Veröffentlicht: (2024)
von: Hou, Zhixing, et al.
Veröffentlicht: (2024)
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
von: Salloom, Tony, et al.
Veröffentlicht: (2025)
von: Salloom, Tony, et al.
Veröffentlicht: (2025)
Geometry-aware 4D Video Generation for Robot Manipulation
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
von: Liu, Zeyi, et al.
Veröffentlicht: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
von: Song, Zijian, et al.
Veröffentlicht: (2026)
von: Song, Zijian, et al.
Veröffentlicht: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
von: Shen, Yichao, et al.
Veröffentlicht: (2025)
RoboDreamer: Learning Compositional World Models for Robot Imagination
von: Zhou, Siyuan, et al.
Veröffentlicht: (2024)
von: Zhou, Siyuan, et al.
Veröffentlicht: (2024)
HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
von: Yuan, Zhecheng, et al.
Veröffentlicht: (2025)
von: Yuan, Zhecheng, et al.
Veröffentlicht: (2025)
RoboPearls: Editable Video Simulation for Robot Manipulation
von: Tang, Tao, et al.
Veröffentlicht: (2025)
von: Tang, Tao, et al.
Veröffentlicht: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
von: Li, Huiqiong, et al.
Veröffentlicht: (2026)
From Nano Robotic Manipulation to Nano Manipulation Robot
von: Zhan Yang, et al.
Veröffentlicht: (2025)
von: Zhan Yang, et al.
Veröffentlicht: (2025)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
von: Zheng, Yuhang, et al.
Veröffentlicht: (2026)
von: Zheng, Yuhang, et al.
Veröffentlicht: (2026)
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
von: Zhang, Kaidong, et al.
Veröffentlicht: (2024)
von: Zhang, Kaidong, et al.
Veröffentlicht: (2024)
Arm Robot: AR-Enhanced Embodied Control and Visualization for Intuitive Robot Arm Manipulation
von: Pei, Siyou, et al.
Veröffentlicht: (2024)
von: Pei, Siyou, et al.
Veröffentlicht: (2024)
MixLight: Borrowing the Best of both Spherical Harmonics and Gaussian Models
von: Ji, Xinlong, et al.
Veröffentlicht: (2024)
von: Ji, Xinlong, et al.
Veröffentlicht: (2024)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
von: Liu, Yifan, et al.
Veröffentlicht: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
von: Bi, Hongzhe, et al.
Veröffentlicht: (2025)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
von: Niu, Yida, et al.
Veröffentlicht: (2026)
von: Niu, Yida, et al.
Veröffentlicht: (2026)
Solving New Tasks by Adapting Internet Video Knowledge
von: Luo, Calvin, et al.
Veröffentlicht: (2025)
von: Luo, Calvin, et al.
Veröffentlicht: (2025)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
von: Chen, Yuzhi, et al.
Veröffentlicht: (2026)
von: Chen, Yuzhi, et al.
Veröffentlicht: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
von: Luo, Yunhao, et al.
Veröffentlicht: (2024)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
von: Guo, Yanjiang, et al.
Veröffentlicht: (2025)
General Neural Gauge Fields
von: Zhan, Fangneng, et al.
Veröffentlicht: (2023)
von: Zhan, Fangneng, et al.
Veröffentlicht: (2023)
DiffAge3D: Diffusion-based 3D-aware Face Aging
von: Wahid, Junaid, et al.
Veröffentlicht: (2024)
von: Wahid, Junaid, et al.
Veröffentlicht: (2024)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahui, et al.
Veröffentlicht: (2025)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
von: Qi, Han, et al.
Veröffentlicht: (2025)
von: Qi, Han, et al.
Veröffentlicht: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
von: Huang, Haifeng, et al.
Veröffentlicht: (2025)
RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning
von: Zhu, Zijian, et al.
Veröffentlicht: (2026)
von: Zhu, Zijian, et al.
Veröffentlicht: (2026)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
von: Wen, Youpeng, et al.
Veröffentlicht: (2024)
von: Wen, Youpeng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
von: Zhou, Kaichen, et al.
Veröffentlicht: (2025) -
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
von: Zhou, Kaichen, et al.
Veröffentlicht: (2026) -
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
von: Liu, Yifan, et al.
Veröffentlicht: (2025) -
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
von: Chang, Xinhai, et al.
Veröffentlicht: (2024) -
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
von: Gu, Renjie, et al.
Veröffentlicht: (2026)