Robot Learning from a Physical World Model
Fuente:
arXiv
Saved in:
| Main Authors: | Mao, Jiageng, He, Sicheng, Wu, Hao-Ning, You, Yang, Sun, Shuyang, Wang, Zhicheng, Bao, Yanan, Chen, Huizhong, Guibas, Leonidas, Guizilini, Vitor, Zhou, Howard, Wang, Yue |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
by: Wu, Yanru, et al.
Published: (2026)
by: Wu, Yanru, et al.
Published: (2026)
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
by: Jia, Emily Yue-Ting, et al.
Published: (2026)
by: Jia, Emily Yue-Ting, et al.
Published: (2026)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)
by: Zhao, Siheng, et al.
Published: (2025)
Humanoid Everyday: A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation
by: Zhao, Zhenyu, et al.
Published: (2025)
by: Zhao, Zhenyu, et al.
Published: (2025)
PhysMem: Scaling Test-Time Memory for Embodied Physical Reasoning
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
Fiducial Exoskeletons: Image-Centric Robot State Estimation
by: Smith, Cameron, et al.
Published: (2026)
by: Smith, Cameron, et al.
Published: (2026)
Towards Realistic Scene Generation with LiDAR Diffusion Models
by: Ran, Haoxi, et al.
Published: (2024)
by: Ran, Haoxi, et al.
Published: (2024)
RoboDream: Compositional World Models for Scalable Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2026)
by: Ye, Junjie, et al.
Published: (2026)
Rodrigues Network for Learning Robot Actions
by: Zhang, Jialiang, et al.
Published: (2025)
by: Zhang, Jialiang, et al.
Published: (2025)
Make a Donut: Hierarchical EMD-Space Planning for Zero-Shot Deformable Manipulation with Tools
by: You, Yang, et al.
Published: (2023)
by: You, Yang, et al.
Published: (2023)
Learning from Massive Human Videos for Universal Humanoid Pose Control
by: Mao, Jiageng, et al.
Published: (2024)
by: Mao, Jiageng, et al.
Published: (2024)
SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment
by: Xue, Rong, et al.
Published: (2025)
by: Xue, Rong, et al.
Published: (2025)
D3RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation
by: Wei, Songlin, et al.
Published: (2024)
by: Wei, Songlin, et al.
Published: (2024)
SparseDFF: Sparse-View Feature Distillation for One-Shot Dexterous Manipulation
by: Wang, Qianxu, et al.
Published: (2023)
by: Wang, Qianxu, et al.
Published: (2023)
AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis
by: Ye, Junjie, et al.
Published: (2025)
by: Ye, Junjie, et al.
Published: (2025)
ARCH: Hierarchical Hybrid Learning for Long-Horizon Contact-Rich Robotic Assembly
by: Sun, Jiankai, et al.
Published: (2024)
by: Sun, Jiankai, et al.
Published: (2024)
PolygMap: A Perceptive Locomotion Framework for Humanoid Robot Stair Climbing
by: Li, Bingquan, et al.
Published: (2025)
by: Li, Bingquan, et al.
Published: (2025)
DOT-Sim: Differentiable Optical Tactile Simulation with Precise Real-to-Sim Physical Calibration
by: You, Yang, et al.
Published: (2026)
by: You, Yang, et al.
Published: (2026)
ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
by: Zhao, Enyu, et al.
Published: (2025)
by: Zhao, Enyu, et al.
Published: (2025)
SAGE: Bridging Semantic and Actionable Parts for GEneralizable Manipulation of Articulated Objects
by: Geng, Haoran, et al.
Published: (2023)
by: Geng, Haoran, et al.
Published: (2023)
Self-Supervised Geometry-Guided Initialization for Robust Monocular Visual Odometry
by: Kanai, Takayuki, et al.
Published: (2024)
by: Kanai, Takayuki, et al.
Published: (2024)
AR Forcing: Towards Long-Horizon Robot Navigation World Model
by: Yang, Yifei, et al.
Published: (2026)
by: Yang, Yifei, et al.
Published: (2026)
Physics-Grounded Differentiable Simulation for Soft Growing Robots
by: Chen, Lucas, et al.
Published: (2025)
by: Chen, Lucas, et al.
Published: (2025)
PhysPart: Physically Plausible Part Completion for Interactable Objects
by: Luo, Rundong, et al.
Published: (2024)
by: Luo, Rundong, et al.
Published: (2024)
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop
by: Hong, Yining, et al.
Published: (2026)
by: Hong, Yining, et al.
Published: (2026)
D-REX: Differentiable Real-to-Sim-to-Real Engine for Learning Dexterous Grasping
by: Lou, Haozhe, et al.
Published: (2026)
by: Lou, Haozhe, et al.
Published: (2026)
A Language Agent for Autonomous Driving
by: Mao, Jiageng, et al.
Published: (2023)
by: Mao, Jiageng, et al.
Published: (2023)
$Ψ_0$: An Open Foundation Model Towards Universal Humanoid Loco-Manipulation
by: Wei, Songlin, et al.
Published: (2026)
by: Wei, Songlin, et al.
Published: (2026)
EquivAct: SIM(3)-Equivariant Visuomotor Policies beyond Rigid Object Manipulation
by: Yang, Jingyun, et al.
Published: (2023)
by: Yang, Jingyun, et al.
Published: (2023)
Kinodynamic Motion Planning for Mobile Robot Navigation across Inconsistent World Models
by: Damm, Eric R., et al.
Published: (2025)
by: Damm, Eric R., et al.
Published: (2025)
Training People to Reward Robots
by: Sun, Endong, et al.
Published: (2025)
by: Sun, Endong, et al.
Published: (2025)
Virtual Community: An Open World for Humans, Robots, and Society
by: Zhou, Qinhong, et al.
Published: (2025)
by: Zhou, Qinhong, et al.
Published: (2025)
Enhancing Diffusion Policy with Classifier-Free Guidance for Temporal Robotic Tasks
by: Lu, Yuang, et al.
Published: (2025)
by: Lu, Yuang, et al.
Published: (2025)
World Models for Robotic Manipulation: A Survey
by: Wang, Fangyuan, et al.
Published: (2026)
by: Wang, Fangyuan, et al.
Published: (2026)
RoboBERT: An End-to-end Multimodal Robotic Manipulation Model
by: Wang, Sicheng, et al.
Published: (2025)
by: Wang, Sicheng, et al.
Published: (2025)
DynaMIC: Dynamic Multimodal In-Context Learning Enabled Embodied Robot Counterfactual Resistance Ability
by: Yan, Tianqiang, et al.
Published: (2025)
by: Yan, Tianqiang, et al.
Published: (2025)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
Subconscious Robotic Imitation Learning
by: Xie, Jun, et al.
Published: (2024)
by: Xie, Jun, et al.
Published: (2024)
Similar Items
-
Large Reward Models: Generalizable Online Robot Reward Generation with Vision-Language Models
by: Wu, Yanru, et al.
Published: (2026) -
DreamPlan: Efficient Reinforcement Fine-Tuning of Vision-Language Planners via Video World Models
by: Jia, Emily Yue-Ting, et al.
Published: (2026) -
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025) -
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024) -
Robot Learning from Any Images
by: Zhao, Siheng, et al.
Published: (2025)