GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Kaichen, Chen, Yuzhen, Zhan, Fangneng, Hua, Hang, Chen, Grace, Chang, Xinhai, Qu, Ao, Du, Yilun, Liu, Zhuang, Liang, Paul Pu, Wang, Mengyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
di: Zhou, Kaichen, et al.
Pubblicazione: (2025)
di: Zhou, Kaichen, et al.
Pubblicazione: (2025)
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
di: Zhou, Kaichen, et al.
Pubblicazione: (2026)
di: Zhou, Kaichen, et al.
Pubblicazione: (2026)
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
di: Liu, Yifan, et al.
Pubblicazione: (2025)
di: Liu, Yifan, et al.
Pubblicazione: (2025)
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
di: Chang, Xinhai, et al.
Pubblicazione: (2024)
di: Chang, Xinhai, et al.
Pubblicazione: (2024)
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
di: Gu, Renjie, et al.
Pubblicazione: (2026)
di: Gu, Renjie, et al.
Pubblicazione: (2026)
Flow Equivariant World Models: Memory for Partially Observed Dynamic Environments
di: Lillemark, Hansen Jin, et al.
Pubblicazione: (2026)
di: Lillemark, Hansen Jin, et al.
Pubblicazione: (2026)
RAD: A Dataset and Benchmark for Real-Life Anomaly Detection with Robotic Observations
di: Zhou, Kaichen, et al.
Pubblicazione: (2024)
di: Zhou, Kaichen, et al.
Pubblicazione: (2024)
Memorization in 3D Shape Generation: An Empirical Study
di: Pu, Shu, et al.
Pubblicazione: (2025)
di: Pu, Shu, et al.
Pubblicazione: (2025)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
di: Zhao, Minda, et al.
Pubblicazione: (2026)
di: Zhao, Minda, et al.
Pubblicazione: (2026)
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
di: Song, Zijian, et al.
Pubblicazione: (2026)
di: Song, Zijian, et al.
Pubblicazione: (2026)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
di: Zhou, Pengfei, et al.
Pubblicazione: (2026)
di: Zhou, Pengfei, et al.
Pubblicazione: (2026)
WorldQA: Multimodal World Knowledge in Videos through Long-Chain Reasoning
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanhan, et al.
Pubblicazione: (2024)
SDP: Spiking Diffusion Policy for Robotic Manipulation with Learnable Channel-Wise Membrane Thresholds
di: Hou, Zhixing, et al.
Pubblicazione: (2024)
di: Hou, Zhixing, et al.
Pubblicazione: (2024)
Geometry-Aware Sparse Depth Sampling for High-Fidelity RGB-D Depth Completion in Robotic Systems
di: Salloom, Tony, et al.
Pubblicazione: (2025)
di: Salloom, Tony, et al.
Pubblicazione: (2025)
Geometry-aware 4D Video Generation for Robot Manipulation
di: Liu, Zeyi, et al.
Pubblicazione: (2025)
di: Liu, Zeyi, et al.
Pubblicazione: (2025)
Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
di: Song, Zijian, et al.
Pubblicazione: (2026)
di: Song, Zijian, et al.
Pubblicazione: (2026)
VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
di: Shen, Yichao, et al.
Pubblicazione: (2025)
di: Shen, Yichao, et al.
Pubblicazione: (2025)
RoboDreamer: Learning Compositional World Models for Robot Imagination
di: Zhou, Siyuan, et al.
Pubblicazione: (2024)
di: Zhou, Siyuan, et al.
Pubblicazione: (2024)
HERMES: Human-to-Robot Embodied Learning from Multi-Source Motion Data for Mobile Dexterous Manipulation
di: Yuan, Zhecheng, et al.
Pubblicazione: (2025)
di: Yuan, Zhecheng, et al.
Pubblicazione: (2025)
RoboPearls: Editable Video Simulation for Robot Manipulation
di: Tang, Tao, et al.
Pubblicazione: (2025)
di: Tang, Tao, et al.
Pubblicazione: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
di: Li, Huiqiong, et al.
Pubblicazione: (2026)
di: Li, Huiqiong, et al.
Pubblicazione: (2026)
From Nano Robotic Manipulation to Nano Manipulation Robot
di: Zhan Yang, et al.
Pubblicazione: (2025)
di: Zhan Yang, et al.
Pubblicazione: (2025)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
di: Zheng, Yuhang, et al.
Pubblicazione: (2026)
di: Zheng, Yuhang, et al.
Pubblicazione: (2026)
PIVOT-R: Primitive-Driven Waypoint-Aware World Model for Robotic Manipulation
di: Zhang, Kaidong, et al.
Pubblicazione: (2024)
di: Zhang, Kaidong, et al.
Pubblicazione: (2024)
Arm Robot: AR-Enhanced Embodied Control and Visualization for Intuitive Robot Arm Manipulation
di: Pei, Siyou, et al.
Pubblicazione: (2024)
di: Pei, Siyou, et al.
Pubblicazione: (2024)
MixLight: Borrowing the Best of both Spherical Harmonics and Gaussian Models
di: Ji, Xinlong, et al.
Pubblicazione: (2024)
di: Ji, Xinlong, et al.
Pubblicazione: (2024)
RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph
di: Liu, Yifan, et al.
Pubblicazione: (2025)
di: Liu, Yifan, et al.
Pubblicazione: (2025)
H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
di: Bi, Hongzhe, et al.
Pubblicazione: (2025)
di: Bi, Hongzhe, et al.
Pubblicazione: (2025)
Scalable Trajectory Generation for Whole-Body Mobile Manipulation
di: Niu, Yida, et al.
Pubblicazione: (2026)
di: Niu, Yida, et al.
Pubblicazione: (2026)
Solving New Tasks by Adapting Internet Video Knowledge
di: Luo, Calvin, et al.
Pubblicazione: (2025)
di: Luo, Calvin, et al.
Pubblicazione: (2025)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
di: Chen, Yuzhi, et al.
Pubblicazione: (2026)
di: Chen, Yuzhi, et al.
Pubblicazione: (2026)
Grounding Video Models to Actions through Goal Conditioned Exploration
di: Luo, Yunhao, et al.
Pubblicazione: (2024)
di: Luo, Yunhao, et al.
Pubblicazione: (2024)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
di: Guo, Yanjiang, et al.
Pubblicazione: (2025)
di: Guo, Yanjiang, et al.
Pubblicazione: (2025)
General Neural Gauge Fields
di: Zhan, Fangneng, et al.
Pubblicazione: (2023)
di: Zhan, Fangneng, et al.
Pubblicazione: (2023)
DiffAge3D: Diffusion-based 3D-aware Face Aging
di: Wahid, Junaid, et al.
Pubblicazione: (2024)
di: Wahid, Junaid, et al.
Pubblicazione: (2024)
SOGS: Second-Order Anchor for Advanced 3D Gaussian Splatting
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
di: Zhang, Jiahui, et al.
Pubblicazione: (2025)
Inference-Time Enhancement of Generative Robot Policies via Predictive World Modeling
di: Qi, Han, et al.
Pubblicazione: (2025)
di: Qi, Han, et al.
Pubblicazione: (2025)
RoboGround: Robotic Manipulation with Grounded Vision-Language Priors
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
di: Huang, Haifeng, et al.
Pubblicazione: (2025)
RDGen: Demonstration Generation for High-Quality Robot Learning via Reinforcement Learning
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
di: Zhu, Zijian, et al.
Pubblicazione: (2026)
VidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation
di: Wen, Youpeng, et al.
Pubblicazione: (2024)
di: Wen, Youpeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
PAGE-4D: VGGT-4D Perception via Disentangled Pose and Geometry Estimation
di: Zhou, Kaichen, et al.
Pubblicazione: (2025) -
Stream3D: Sequential Multi-View 3D Generation via Evidential Memory
di: Zhou, Kaichen, et al.
Pubblicazione: (2026) -
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
di: Liu, Yifan, et al.
Pubblicazione: (2025) -
Neural Surface Reconstruction from Sparse Views Using Epipolar Geometry
di: Chang, Xinhai, et al.
Pubblicazione: (2024) -
GeoWorld-VLM: Geometry from World Models for Vision-Language Models
di: Gu, Renjie, et al.
Pubblicazione: (2026)