World Model for Robot Learning: A Comprehensive Survey
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Bohan, Li, Gen, Jia, Jindou, An, Tuo, Guo, Xinying, Leng, Sicong, Geng, Haoran, Ze, Yanjie, Harada, Tatsuya, Torr, Philip, Mees, Oier, Pollefeys, Marc, Liu, Zhuang, Wu, Jiajun, Abbeel, Pieter, Malik, Jitendra, Du, Yilun, Yang, Jianfei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
by: Tang, Yikai, et al.
Published: (2025)
by: Tang, Yikai, et al.
Published: (2025)
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
by: Heng, Liang, et al.
Published: (2025)
by: Heng, Liang, et al.
Published: (2025)
Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
by: Jones, Joshua, et al.
Published: (2025)
by: Jones, Joshua, et al.
Published: (2025)
Action-to-Action Flow Matching
by: Jia, Jindou, et al.
Published: (2026)
by: Jia, Jindou, et al.
Published: (2026)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Interactive Task Planning with Language Models
by: Li, Boyi, et al.
Published: (2023)
by: Li, Boyi, et al.
Published: (2023)
Deep Sensorimotor Control by Imitating Predictive Models of Human Motion
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
by: Singh, Himanshu Gaurav, et al.
Published: (2025)
Multi-Objective Learning for Diffusion Models: A Statistical Theory under Semi-Supervised Learning
by: Cheng, Ziheng, et al.
Published: (2026)
by: Cheng, Ziheng, et al.
Published: (2026)
Feedback World Model Enables Precise Guidance of Diffusion Policy
by: An, Tuo, et al.
Published: (2026)
by: An, Tuo, et al.
Published: (2026)
SkillBlender: Towards Versatile Humanoid Whole-Body Loco-Manipulation via Skill Blending
by: Kuang, Yuxuan, et al.
Published: (2025)
by: Kuang, Yuxuan, et al.
Published: (2025)
How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
by: Lin, Toru, et al.
Published: (2026)
by: Lin, Toru, et al.
Published: (2026)
Twisting Lids Off with Two Hands
by: Lin, Toru, et al.
Published: (2024)
by: Lin, Toru, et al.
Published: (2024)
MARS Policy: Multimodality Only When It Matters
by: Jia, Jindou, et al.
Published: (2026)
by: Jia, Jindou, et al.
Published: (2026)
Large Video Planner Enables Generalizable Robot Control
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
FLASH: Efficient Visuomotor Policy via Sparse Sampling
by: Bai, Jiaqi, et al.
Published: (2026)
by: Bai, Jiaqi, et al.
Published: (2026)
DexGarmentLab: Dexterous Garment Manipulation Environment with Generalizable Policy
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
End-to-end RL Improves Dexterous Grasping Policies
by: Singh, Ritvik, et al.
Published: (2025)
by: Singh, Ritvik, et al.
Published: (2025)
The Sound of Simulation: Learning Multimodal Sim-to-Real Robot Policies with Generative Audio
by: Wang, Renhao, et al.
Published: (2025)
by: Wang, Renhao, et al.
Published: (2025)
RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark
by: Cui, Ziteng, et al.
Published: (2025)
by: Cui, Ziteng, et al.
Published: (2025)
Steering Your Generalists: Improving Robotic Foundation Models via Value Guidance
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
by: Nakamoto, Mitsuhiko, et al.
Published: (2024)
Multimodal Spatial Language Maps for Robot Navigation and Manipulation
by: Huang, Chenguang, et al.
Published: (2025)
by: Huang, Chenguang, et al.
Published: (2025)
Vision-Language Models Provide Promptable Representations for Reinforcement Learning
by: Chen, William, et al.
Published: (2024)
by: Chen, William, et al.
Published: (2024)
TWIST2: Scalable, Portable, and Holistic Humanoid Data Collection System
by: Ze, Yanjie, et al.
Published: (2025)
by: Ze, Yanjie, et al.
Published: (2025)
Hand-Object Interaction Pretraining from Videos
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
by: Singh, Himanshu Gaurav, et al.
Published: (2024)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
by: Lou, Haozhe, et al.
Published: (2026)
by: Lou, Haozhe, et al.
Published: (2026)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
by: Li, Jingliang, et al.
Published: (2026)
by: Li, Jingliang, et al.
Published: (2026)
ResMimic: From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
OTTER: A Vision-Language-Action Model with Text-Aware Visual Feature Extraction
by: Huang, Huang, et al.
Published: (2025)
by: Huang, Huang, et al.
Published: (2025)
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
by: Hou, Bohan, et al.
Published: (2026)
by: Hou, Bohan, et al.
Published: (2026)
FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control
by: Seo, Younggyo, et al.
Published: (2025)
by: Seo, Younggyo, et al.
Published: (2025)
Coarse-to-fine Q-Network with Action Sequence for Data-Efficient Reinforcement Learning
by: Seo, Younggyo, et al.
Published: (2024)
by: Seo, Younggyo, et al.
Published: (2024)
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation
by: Doshi, Ria, et al.
Published: (2024)
by: Doshi, Ria, et al.
Published: (2024)
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images
by: Chen, Yuedong, et al.
Published: (2024)
by: Chen, Yuedong, et al.
Published: (2024)
Policy Adaptation via Language Optimization: Decomposing Tasks for Few-Shot Imitation
by: Myers, Vivek, et al.
Published: (2024)
by: Myers, Vivek, et al.
Published: (2024)
The Ingredients for Robotic Diffusion Transformers
by: Dasari, Sudeep, et al.
Published: (2024)
by: Dasari, Sudeep, et al.
Published: (2024)
Lightning Grasp: High Performance Procedural Grasp Synthesis with Contact Fields
by: Yin, Zhao-Heng, et al.
Published: (2025)
by: Yin, Zhao-Heng, et al.
Published: (2025)
Offline Imitation Learning Through Graph Search and Retrieval
by: Yin, Zhao-Heng, et al.
Published: (2024)
by: Yin, Zhao-Heng, et al.
Published: (2024)
X-Capture: An Open-Source Portable Device for Multi-Sensory Learning
by: Clarke, Samuel, et al.
Published: (2025)
by: Clarke, Samuel, et al.
Published: (2025)
Similar Items
-
DIPOLE: Fusing Vision and Geometry for Robust Visuomotor Generalization
by: Tang, Yikai, et al.
Published: (2025) -
ViTacFormer: Learning Cross-Modal Representation for Visuo-Tactile Dexterous Manipulation
by: Heng, Liang, et al.
Published: (2025) -
Beyond Sight: Finetuning Generalist Robot Policies with Heterogeneous Sensors via Language Grounding
by: Jones, Joshua, et al.
Published: (2025) -
Action-to-Action Flow Matching
by: Jia, Jindou, et al.
Published: (2026) -
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)