Disentangled Robot Learning via Separate Forward and Inverse Dynamics Pretraining
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wenyao, Zhang, Bozhou, Qi, Zekun, Zeng, Wenjun, Jin, Xin, Zhang, Li |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
by: Peng, Baorui, et al.
Published: (2026)
by: Peng, Baorui, et al.
Published: (2026)
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
by: Xu, Tianyu, et al.
Published: (2025)
by: Xu, Tianyu, et al.
Published: (2025)
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States
by: Zhang, Bozhou, et al.
Published: (2024)
by: Zhang, Bozhou, et al.
Published: (2024)
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
VLA-JEPA: Enhancing Vision-Language-Action Model with Latent World Model
by: Sun, Jingwen, et al.
Published: (2026)
by: Sun, Jingwen, et al.
Published: (2026)
DexVLG: Dexterous Vision-Language-Grasp Model at Scale
by: He, Jiawei, et al.
Published: (2025)
by: He, Jiawei, et al.
Published: (2025)
From Imagined Futures to Executable Actions: Mixture of Latent Actions for Robot Manipulation
by: Li, Yajie, et al.
Published: (2026)
by: Li, Yajie, et al.
Published: (2026)
Relative Position Matters: Trajectory Prediction and Planning with Polar Representation
by: Zhang, Bozhou, et al.
Published: (2025)
by: Zhang, Bozhou, et al.
Published: (2025)
Reasoning in Space via Grounding in the World
by: Chen, Yiming, et al.
Published: (2025)
by: Chen, Yiming, et al.
Published: (2025)
Scene Graph Disentanglement and Composition for Generalizable Complex Image Generation
by: Wang, Yunnan, et al.
Published: (2024)
by: Wang, Yunnan, et al.
Published: (2024)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
STORM: Search-Guided Generative World Models for Robotic Manipulation
by: Lin, Wenjun, et al.
Published: (2025)
by: Lin, Wenjun, et al.
Published: (2025)
SpatialActor: Exploring Disentangled Spatial Representations for Robust Robotic Manipulation
by: Shi, Hao, et al.
Published: (2025)
by: Shi, Hao, et al.
Published: (2025)
VLS: Steering Pretrained Robot Policies via Vision-Language Models
by: Liu, Shuo, et al.
Published: (2026)
by: Liu, Shuo, et al.
Published: (2026)
Closed-Loop Unsupervised Representation Disentanglement with $β$-VAE Distillation and Diffusion Probabilistic Feedback
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
ManiVID-3D: Generalizable View-Invariant Reinforcement Learning for Robotic Manipulation via Disentangled 3D Representations
by: Li, Zheng, et al.
Published: (2025)
by: Li, Zheng, et al.
Published: (2025)
SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object Manipulation
by: Qi, Zekun, et al.
Published: (2025)
by: Qi, Zekun, et al.
Published: (2025)
GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
by: Pei, Muleilan, et al.
Published: (2025)
by: Pei, Muleilan, et al.
Published: (2025)
Don't Let Your Robot be Harmful: Responsible Robotic Manipulation via Safety-as-Policy
by: Ni, Minheng, et al.
Published: (2024)
by: Ni, Minheng, et al.
Published: (2024)
Hierarchical Temporal Context Learning for Camera-based Semantic Scene Completion
by: Li, Bohan, et al.
Published: (2024)
by: Li, Bohan, et al.
Published: (2024)
SpectralSplat: Appearance-Disentangled Feed-Forward Gaussian Splatting for Driving Scenes
by: Herau, Quentin, et al.
Published: (2026)
by: Herau, Quentin, et al.
Published: (2026)
Hybrid-grained Feature Aggregation with Coarse-to-fine Language Guidance for Self-supervised Monocular Depth Estimation
by: Zhang, Wenyao, et al.
Published: (2025)
by: Zhang, Wenyao, et al.
Published: (2025)
Disentangled Object-Centric Image Representation for Robotic Manipulation
by: Emukpere, David, et al.
Published: (2025)
by: Emukpere, David, et al.
Published: (2025)
Decision-Driven Semantic Object Exploration for Legged Robots via Confidence-Calibrated Perception and Topological Subgoal Selection
by: Zhao, Guoyang, et al.
Published: (2025)
by: Zhao, Guoyang, et al.
Published: (2025)
PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
by: Wan, Jiansong, et al.
Published: (2025)
by: Wan, Jiansong, et al.
Published: (2025)
MapGS: Generalizable Pretraining and Data Augmentation for Online Mapping via Novel View Synthesis
by: Zhang, Hengyuan, et al.
Published: (2025)
by: Zhang, Hengyuan, et al.
Published: (2025)
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
by: Wen, Xin, et al.
Published: (2025)
by: Wen, Xin, et al.
Published: (2025)
DDP-WM: Disentangled Dynamics Prediction for Efficient World Models
by: Yin, Shicheng, et al.
Published: (2026)
by: Yin, Shicheng, et al.
Published: (2026)
Robo-MUTUAL: Robotic Multimodal Task Specification via Unimodal Learning
by: Li, Jianxiong, et al.
Published: (2024)
by: Li, Jianxiong, et al.
Published: (2024)
Efficient Long-Horizon Vision-Language-Action Models via Static-Dynamic Disentanglement
by: Qiu, Weikang, et al.
Published: (2026)
by: Qiu, Weikang, et al.
Published: (2026)
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
ClearDepth: Enhanced Stereo Perception of Transparent Objects for Robotic Manipulation
by: Bai, Kaixin, et al.
Published: (2024)
by: Bai, Kaixin, et al.
Published: (2024)
PhysReaction: Physically Plausible Real-Time Humanoid Reaction Synthesis via Forward Dynamics Guided 4D Imitation
by: Liu, Yunze, et al.
Published: (2024)
by: Liu, Yunze, et al.
Published: (2024)
Interpretable Single-View 3D Gaussian Splatting using Unsupervised Hierarchical Disentangled Representation Learning
by: Zhang, Yuyang, et al.
Published: (2025)
by: Zhang, Yuyang, et al.
Published: (2025)
GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation
by: Chai, Ying, et al.
Published: (2025)
by: Chai, Ying, et al.
Published: (2025)
Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation
by: Almuzairee, Abdulaziz, et al.
Published: (2025)
by: Almuzairee, Abdulaziz, et al.
Published: (2025)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
Similar Items
-
ReWorld: Multi-Dimensional Reward Modeling for Embodied World Models
by: Peng, Baorui, et al.
Published: (2026) -
MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
by: Xu, Tianyu, et al.
Published: (2025) -
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
by: Zhang, Bozhou, et al.
Published: (2025) -
DeMo: Decoupling Motion Forecasting into Directional Intentions and Dynamic States
by: Zhang, Bozhou, et al.
Published: (2024) -
DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
by: Zhang, Wenyao, et al.
Published: (2025)