GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhou, Kaichen, Chen, Yuzhen, Zhan, Fangneng, Hua, Hang, Chen, Grace, Chang, Xinhai, Qu, Ao, Du, Yilun, Liu, Zhuang, Liang, Paul Pu, Wang, Mengyu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
by: Zhou, Kaichen, et al.
Published: (2025)
by: Zhou, Kaichen, et al.
Published: (2025)
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
by: Chang, Xinhai, et al.
Published: (2024)
by: Chang, Xinhai, et al.
Published: (2024)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)
by: Gu, Renjie, et al.
Published: (2026)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
by: Lillemark, Hansen Jin, et al.
Published: (2026)
by: Lillemark, Hansen Jin, et al.
Published: (2026)
RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations
by: Zhou, Kaichen, et al.
Published: (2024)
by: Zhou, Kaichen, et al.
Published: (2024)
Memorization in 3D Shape Generation: An Empirical Study
by: Pu, Shu, et al.
Published: (2025)
by: Pu, Shu, et al.
Published: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
by: Zhao, Minda, et al.
Published: (2026)
by: Zhao, Minda, et al.
Published: (2026)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
by: Zhou, Pengfei, et al.
Published: (2026)
by: Zhou, Pengfei, et al.
Published: (2026)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
by: Zhang, Yuanhan, et al.
Published: (2024)
by: Zhang, Yuanhan, et al.
Published: (2024)
SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds
by: Hou, Zhixing, et al.
Published: (2024)
by: Hou, Zhixing, et al.
Published: (2024)
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
by: Salloom, Tony, et al.
Published: (2025)
by: Salloom, Tony, et al.
Published: (2025)
Geometry-aware 4D Video Generation for Robot Manipulation
by: Liu, Zeyi, et al.
Published: (2025)
by: Liu, Zeyi, et al.
Published: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
by: Shen, Yichao, et al.
Published: (2025)
by: Shen, Yichao, et al.
Published: (2025)
RoboDreamer: Learning Compositional World Models for Robot Imagination
by: Zhou, Siyuan, et al.
Published: (2024)
by: Zhou, Siyuan, et al.
Published: (2024)
HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
by: Yuan, Zhecheng, et al.
Published: (2025)
by: Yuan, Zhecheng, et al.
Published: (2025)
RoboPearls: Editable Video Simulation for Robot Manipulation
by: Tang, Tao, et al.
Published: (2025)
by: Tang, Tao, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
From Nano Robotic Manipulation to Nano Manipulation Robot
by: Zhan Yang, et al.
Published: (2025)
by: Zhan Yang, et al.
Published: (2025)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026)
by: Zheng, Yuhang, et al.
Published: (2026)
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
by: Zhang, Kaidong, et al.
Published: (2024)
by: Zhang, Kaidong, et al.
Published: (2024)
Arm Robot: AR-Enhanced Embodied Control and Visualization for Intuitive Robot Arm Manipulation
by: Pei, Siyou, et al.
Published: (2024)
by: Pei, Siyou, et al.
Published: (2024)
MixLight: Borrowing the Best of both Spherical Harmonics and Gaussian Models
by: Ji, Xinlong, et al.
Published: (2024)
by: Ji, Xinlong, et al.
Published: (2024)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
by: Liu, Yifan, et al.
Published: (2025)
by: Liu, Yifan, et al.
Published: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
by: Bi, Hongzhe, et al.
Published: (2025)
by: Bi, Hongzhe, et al.
Published: (2025)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
by: Niu, Yida, et al.
Published: (2026)
by: Niu, Yida, et al.
Published: (2026)
Solving New Tasks by Adapting Internet Video Knowledge
by: Luo, Calvin, et al.
Published: (2025)
by: Luo, Calvin, et al.
Published: (2025)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026)
by: Chen, Yuzhi, et al.
Published: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
by: Luo, Yunhao, et al.
Published: (2024)
by: Luo, Yunhao, et al.
Published: (2024)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
General Neural Gauge Fields
by: Zhan, Fangneng, et al.
Published: (2023)
by: Zhan, Fangneng, et al.
Published: (2023)
DiffAge3D: Diffusion-based 3D-aware Face Aging
by: Wahid, Junaid, et al.
Published: (2024)
by: Wahid, Junaid, et al.
Published: (2024)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
by: Zhang, Jiahui, et al.
Published: (2025)
by: Zhang, Jiahui, et al.
Published: (2025)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
by: Qi, Han, et al.
Published: (2025)
by: Qi, Han, et al.
Published: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
by: Huang, Haifeng, et al.
Published: (2025)
by: Huang, Haifeng, et al.
Published: (2025)
RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning
by: Zhu, Zijian, et al.
Published: (2026)
by: Zhu, Zijian, et al.
Published: (2026)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
by: Wen, Youpeng, et al.
Published: (2024)
by: Wen, Youpeng, et al.
Published: (2024)
Similar Items
-
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
by: Zhou, Kaichen, et al.
Published: (2025) -
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
by: Zhou, Kaichen, et al.
Published: (2026) -
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
by: Liu, Yifan, et al.
Published: (2025) -
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
by: Chang, Xinhai, et al.
Published: (2024) -
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
by: Gu, Renjie, et al.
Published: (2026)