Learning Physics from Pretrained Video Models: A Multimodal Continuous and Sequential World Interaction Models for Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Song, Zijian, Li, Qichang, Qin, Sihan, Chen, Yuhao, Chen, Tianshui, Lin, Liang, Wang, Guangrun |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
by: Song, Zijian, et al.
Published: (2026)
by: Song, Zijian, et al.
Published: (2026)
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025)
by: Song, Zijian, et al.
Published: (2025)
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
by: Chen, Yuhao, et al.
Published: (2026)
by: Chen, Yuhao, et al.
Published: (2026)
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026)
by: Chen, Yuzhi, et al.
Published: (2026)
Hyperbolic Multiview Pretraining for Robotic Manipulation
by: Yang, Jin, et al.
Published: (2026)
by: Yang, Jin, et al.
Published: (2026)
Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
by: Li, Qixiu, et al.
Published: (2025)
by: Li, Qixiu, et al.
Published: (2025)
$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
by: Zhou, Pengfei, et al.
Published: (2026)
by: Zhou, Pengfei, et al.
Published: (2026)
GEM-4D: Geometry-Enhanced Video World Models for Robot Manipulation
by: Zhou, Kaichen, et al.
Published: (2026)
by: Zhou, Kaichen, et al.
Published: (2026)
ManipulationNet: An Infrastructure for Benchmarking Real-World Robot Manipulation with Physical Skill Challenges and Embodied Multimodal Reasoning
by: Chen, Yiting, et al.
Published: (2026)
by: Chen, Yiting, et al.
Published: (2026)
Say, Dream, and Act: Learning Video World Models for Instruction-Driven Robot Manipulation
by: Gu, Songen, et al.
Published: (2026)
by: Gu, Songen, et al.
Published: (2026)
VLA Models Are More Generalizable Than You Think: Revisiting Physical and Spatial Modeling
by: Li, Weiqi, et al.
Published: (2025)
by: Li, Weiqi, et al.
Published: (2025)
World Models for Robotic Manipulation: A Survey
by: Wang, Fangyuan, et al.
Published: (2026)
by: Wang, Fangyuan, et al.
Published: (2026)
Stable Language Guidance for Vision-Language-Action Models
by: Zhan, Zhihao, et al.
Published: (2026)
by: Zhan, Zhihao, et al.
Published: (2026)
Learning Robot Manipulation from Audio World Models
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
Empowering Multi-Robot Cooperation via Sequential World Models
by: Zhao, Zijie, et al.
Published: (2025)
by: Zhao, Zijie, et al.
Published: (2025)
World4RL: Diffusion World Models for Policy Refinement with Reinforcement Learning for Robotic Manipulation
by: Jiang, Zhennan, et al.
Published: (2025)
by: Jiang, Zhennan, et al.
Published: (2025)
RoboTrustBench: Benchmarking the Trustworthiness of Video World Models for Robotic Manipulation
by: Li, Huiqiong, et al.
Published: (2026)
by: Li, Huiqiong, et al.
Published: (2026)
GaussianDream: A Feed-Forward 3D Gaussian World Model for Robotic Manipulation
by: Zhang, Zijian, et al.
Published: (2026)
by: Zhang, Zijian, et al.
Published: (2026)
OmniVTA: Visuo-Tactile World Modeling for Contact-Rich Robotic Manipulation
by: Zheng, Yuhang, et al.
Published: (2026)
by: Zheng, Yuhang, et al.
Published: (2026)
Ctrl-World: A Controllable Generative World Model for Robot Manipulation
by: Guo, Yanjiang, et al.
Published: (2025)
by: Guo, Yanjiang, et al.
Published: (2025)
H2R: A Human-to-Robot Data Augmentation for Robot Pre-training from Videos
by: Li, Guangrun, et al.
Published: (2025)
by: Li, Guangrun, et al.
Published: (2025)
Visual Robotic Manipulation with Depth-Aware Pretraining
by: Wang, Wanying, et al.
Published: (2024)
by: Wang, Wanying, et al.
Published: (2024)
iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation
by: Zhang, Chuanrui, et al.
Published: (2025)
by: Zhang, Chuanrui, et al.
Published: (2025)
PIN-WM: Learning Physics-INformed World Models for Non-Prehensile Manipulation
by: Li, Wenxuan, et al.
Published: (2025)
by: Li, Wenxuan, et al.
Published: (2025)
RoboHorizon: An LLM-Assisted Multi-View World Model for Long-Horizon Robotic Manipulation
by: Chen, Zixuan, et al.
Published: (2025)
by: Chen, Zixuan, et al.
Published: (2025)
Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction
by: Jiang, Shuo, et al.
Published: (2025)
by: Jiang, Shuo, et al.
Published: (2025)
Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
by: Bai, Shuanghao, et al.
Published: (2025)
by: Bai, Shuanghao, et al.
Published: (2025)
STARRY: Spatial-Temporal Action-Centric World Modeling for Robotic Manipulation
by: Tian, Yuxuan, et al.
Published: (2026)
by: Tian, Yuxuan, et al.
Published: (2026)
OA-WAM: Object-Addressable World Action Model for Robust Robot Manipulation
by: Liu, Yushan, et al.
Published: (2026)
by: Liu, Yushan, et al.
Published: (2026)
A Recipe for Efficient Sim-to-Real Transfer in Manipulation with Online Imitation-Pretrained World Models
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
A General One-Shot Multimodal Active Perception Framework for Robotic Manipulation: Learning to Predict Optimal Viewpoint
by: Qin, Deyun, et al.
Published: (2026)
by: Qin, Deyun, et al.
Published: (2026)
Mastering Robot Manipulation with Multimodal Prompts through Pretraining and Multi-task Fine-tuning
by: Li, Jiachen, et al.
Published: (2023)
by: Li, Jiachen, et al.
Published: (2023)
What Foundation Models can Bring for Robot Learning in Manipulation : A Survey
by: Li, Dingzhe, et al.
Published: (2024)
by: Li, Dingzhe, et al.
Published: (2024)
Dynamic Policy Learning for Legged Robot with Simplified Model Pretraining and Model Homotopy Transfer
by: Kang, Dongyun, et al.
Published: (2025)
by: Kang, Dongyun, et al.
Published: (2025)
Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
by: Zhu, Chuning, et al.
Published: (2025)
by: Zhu, Chuning, et al.
Published: (2025)
DexSim2Real: Foundation Model-Guided Sim-to-Real Transfer for Generalizable Dexterous Manipulation
by: Zeng, Zijian, et al.
Published: (2026)
by: Zeng, Zijian, et al.
Published: (2026)
RoTri-Diff: A Spatial Robot-Object Triadic Interaction-Guided Diffusion Model for Bimanual Manipulation
by: Chen, Zixuan, et al.
Published: (2026)
by: Chen, Zixuan, et al.
Published: (2026)
RISE: Self-Improving Robot Policy with Compositional World Model
by: Yang, Jiazhi, et al.
Published: (2026)
by: Yang, Jiazhi, et al.
Published: (2026)
Building Explicit World Model for Zero-Shot Open-World Object Manipulation
by: Li, Xiaotong, et al.
Published: (2026)
by: Li, Xiaotong, et al.
Published: (2026)
DeformMaster: An Interactive Physics-Neural World Model for Deformable Objects from Videos
by: Li, Can, et al.
Published: (2026)
by: Li, Can, et al.
Published: (2026)
Similar Items
-
Robotic Manipulation is Vision-to-Geometry Mapping ($f(v) \rightarrow G$): Vision-Geometry Backbones over Language and Video Models
by: Song, Zijian, et al.
Published: (2026) -
Physical Autoregressive Model for Robotic Manipulation without Action Pretraining
by: Song, Zijian, et al.
Published: (2025) -
RADAR: Benchmarking Vision-Language-Action Generalization via Real-World Dynamics, Spatial-Physical Intelligence, and Autonomous Evaluation
by: Chen, Yuhao, et al.
Published: (2026) -
ABot-PhysWorld: Interactive World Foundation Model for Robotic Manipulation with Physics Alignment
by: Chen, Yuzhi, et al.
Published: (2026) -
Hyperbolic Multiview Pretraining for Robotic Manipulation
by: Yang, Jin, et al.
Published: (2026)