CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Jiange, Shi, Yansong, Zhu, Haoyi, Liu, Mingyu, Ma, Kaijing, Wang, Yating, Wu, Gangshan, He, Tong, Wang, Limin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning
by: Yang, Jiange, et al.
Published: (2024)
by: Yang, Jiange, et al.
Published: (2024)
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025)
by: Sun, Qiao, et al.
Published: (2025)
Spatiotemporal Predictive Pre-training for Robotic Motor Control
by: Yang, Jiange, et al.
Published: (2024)
by: Yang, Jiange, et al.
Published: (2024)
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)
by: Wang, Yating, et al.
Published: (2025)
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
by: Liu, Mingyu, et al.
Published: (2025)
by: Liu, Mingyu, et al.
Published: (2025)
Transferring Foundation Models for Generalizable Robotic Manipulation
by: Yang, Jiange, et al.
Published: (2023)
by: Yang, Jiange, et al.
Published: (2023)
Point Cloud Matters: Rethinking the Impact of Different Observation Spaces on Robot Learning
by: Zhu, Haoyi, et al.
Published: (2024)
by: Zhu, Haoyi, et al.
Published: (2024)
CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems
by: Cao, Zhihao, et al.
Published: (2026)
by: Cao, Zhihao, et al.
Published: (2026)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
by: Chen, Yi, et al.
Published: (2024)
by: Chen, Yi, et al.
Published: (2024)
CoMo: Compositional Motion Customization for Text-to-Video Generation
by: Xu, Youcan, et al.
Published: (2025)
by: Xu, Youcan, et al.
Published: (2025)
Deadlock-Free Hybrid RL-MAPF Framework for Zero-Shot Multi-Robot Navigation
by: Wang, Haoyi, et al.
Published: (2025)
by: Wang, Haoyi, et al.
Published: (2025)
VLA-RL: Towards Masterful and General Robotic Manipulation with Scalable Reinforcement Learning
by: Lu, Guanxing, et al.
Published: (2025)
by: Lu, Guanxing, et al.
Published: (2025)
ConLA: Contrastive Latent Action Learning from Human Videos for Robotic Manipulation
by: Dai, Weisheng, et al.
Published: (2026)
by: Dai, Weisheng, et al.
Published: (2026)
iMoWM: Taming Interactive Multi-Modal World Model for Robotic Manipulation
by: Zhang, Chuanrui, et al.
Published: (2025)
by: Zhang, Chuanrui, et al.
Published: (2025)
Contrast, Imitate, Adapt: Learning Robotic Skills From Raw Human Videos
by: Qian, Zhifeng, et al.
Published: (2024)
by: Qian, Zhifeng, et al.
Published: (2024)
Learning Action Manifold with Multi-view Latent Priors for Robotic Manipulation
by: Xiao, Junjin, et al.
Published: (2026)
by: Xiao, Junjin, et al.
Published: (2026)
LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning
by: Rodriguez, Ariel, et al.
Published: (2026)
by: Rodriguez, Ariel, et al.
Published: (2026)
Efficient Robotic Policy Learning via Latent Space Backward Planning
by: Liu, Dongxiu, et al.
Published: (2025)
by: Liu, Dongxiu, et al.
Published: (2025)
MemoAct: Atkinson-Shiffrin-Inspired Memory-Augmented Visuomotor Policy for Robotic Manipulation
by: Tan, Liufan, et al.
Published: (2026)
by: Tan, Liufan, et al.
Published: (2026)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
CoLA-Flow Policy: Temporally Coherent Imitation Learning via Continuous Latent Action Flow Matching for Robotic Manipulation
by: Songwei, Wu, et al.
Published: (2026)
by: Songwei, Wu, et al.
Published: (2026)
Bridging Deep Reinforcement Learning and Motion Planning for Model-Free Navigation in Cluttered Environments
by: Luo, Licheng, et al.
Published: (2025)
by: Luo, Licheng, et al.
Published: (2025)
CORAL: Scalable Multi-Task Robot Learning via LoRA Experts
by: Luo, Yuankai, et al.
Published: (2026)
by: Luo, Yuankai, et al.
Published: (2026)
CLAP: Contrastive Latent Action Pretraining for Learning Vision-Language-Action Models from Human Videos
by: Zhang, Chubin, et al.
Published: (2026)
by: Zhang, Chubin, et al.
Published: (2026)
RoMoCo: Robotic Motion Control Toolbox for Reduced-Order Model-Based Locomotion on Bipedal and Humanoid Robots
by: Dai, Min, et al.
Published: (2025)
by: Dai, Min, et al.
Published: (2025)
MA-ROESL: Motion-aware Rapid Reward Optimization for Efficient Robot Skill Learning from Single Videos
by: Wang, Xianghui, et al.
Published: (2025)
by: Wang, Xianghui, et al.
Published: (2025)
Towards Generalist Robot Learning from Internet Video: A Survey
by: McCarthy, Robert, et al.
Published: (2024)
by: McCarthy, Robert, et al.
Published: (2024)
Sensorimotor Attention and Language-based Regressions in Shared Latent Variables for Integrating Robot Motion Learning and LLM
by: Suzuki, Kanata, et al.
Published: (2024)
by: Suzuki, Kanata, et al.
Published: (2024)
ATRos: Learning Energy-Efficient Agile Locomotion for Wheeled-legged Robots
by: Sun, Jingyuan, et al.
Published: (2025)
by: Sun, Jingyuan, et al.
Published: (2025)
Learning High-Frequency Continuous Action Chunks in Latent Space
by: Wang, Kunyun, et al.
Published: (2026)
by: Wang, Kunyun, et al.
Published: (2026)
Distributed Non-Uniform Scaling Control of Multi-Agent Formation with Dynamic Agent Joining
by: He, Tao, et al.
Published: (2026)
by: He, Tao, et al.
Published: (2026)
Scalable Dexterous Robot Learning with AR-based Remote Human-Robot Interactions
by: Yang, Yicheng, et al.
Published: (2026)
by: Yang, Yicheng, et al.
Published: (2026)
Video-to-BT: Generating Reactive Behavior Trees from Human Demonstration Videos for Robotic Assembly
by: Zhao, Xiwei, et al.
Published: (2025)
by: Zhao, Xiwei, et al.
Published: (2025)
Learn Weightlessness: Imitate Non-Self-Stabilizing Motions on Humanoid Robot
by: Xin, Yucheng, et al.
Published: (2026)
by: Xin, Yucheng, et al.
Published: (2026)
CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control
by: Shi, Jiaqi, et al.
Published: (2026)
by: Shi, Jiaqi, et al.
Published: (2026)
Genie Centurion: Accelerating Scalable Real-World Robot Training with Human Rewind-and-Refine Guidance
by: Wang, Wenhao, et al.
Published: (2025)
by: Wang, Wenhao, et al.
Published: (2025)
Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids
by: Hu, Kaizhe, et al.
Published: (2025)
by: Hu, Kaizhe, et al.
Published: (2025)
Robot Learning from Human Videos: A Survey
by: Ma, Junyi, et al.
Published: (2026)
by: Ma, Junyi, et al.
Published: (2026)
Hybrid Robot Learning for Automatic Robot Motion Planning in Manufacturing
by: Singh, Siddharth, et al.
Published: (2025)
by: Singh, Siddharth, et al.
Published: (2025)
Similar Items
-
Tra-MoE: Learning Trajectory Prediction Model from Multiple Domains for Adaptive Policy Conditioning
by: Yang, Jiange, et al.
Published: (2024) -
SPA: 3D Spatial-Awareness Enables Effective Embodied Representation
by: Zhu, Haoyi, et al.
Published: (2024) -
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
by: Sun, Qiao, et al.
Published: (2025) -
Spatiotemporal Predictive Pre-training for Robotic Motor Control
by: Yang, Jiange, et al.
Published: (2024) -
VQ-VLA: Improving Vision-Language-Action Models via Scaling Vector-Quantized Action Tokenizers
by: Wang, Yating, et al.
Published: (2025)