MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhang, Qiang, Ma, Jiahao, Liu, Peiran, Shi, Shuai, Su, Zeran, Wang, Zifan, Sun, Jingkai, Cui, Wei, Yu, Jialin, Han, Gang, Zhao, Wen, Sun, Pihai, Yin, Kangning, Wang, Jiaxu, Cao, Jiahang, Zhang, Lingfeng, Cheng, Hao, Hao, Xiaoshuai, Ji, Yiding, Liang, Junwei, Tang, Jian, Xu, Renjing, Guo, Yijie
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910024963457024
author Zhang, Qiang
Ma, Jiahao
Liu, Peiran
Shi, Shuai
Su, Zeran
Wang, Zifan
Sun, Jingkai
Cui, Wei
Yu, Jialin
Han, Gang
Zhao, Wen
Sun, Pihai
Yin, Kangning
Wang, Jiaxu
Cao, Jiahang
Zhang, Lingfeng
Cheng, Hao
Hao, Xiaoshuai
Ji, Yiding
Liang, Junwei
Tang, Jian
Xu, Renjing
Guo, Yijie
author_facet Zhang, Qiang
Ma, Jiahao
Liu, Peiran
Shi, Shuai
Su, Zeran
Wang, Zifan
Sun, Jingkai
Cui, Wei
Yu, Jialin
Han, Gang
Zhao, Wen
Sun, Pihai
Yin, Kangning
Wang, Jiaxu
Cao, Jiahang
Zhang, Lingfeng
Cheng, Hao
Hao, Xiaoshuai
Ji, Yiding
Liang, Junwei
Tang, Jian
Xu, Renjing
Guo, Yijie
contents Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and intricate dynamics of humanoid robots make manual motion design impractical, leading to a heavy reliance on expensive motion capture (MoCap) data. These datasets are not only costly to acquire but also frequently lack the necessary geometric context of the surrounding physical environment. Consequently, existing motion synthesis frameworks often suffer from a decoupling of motion and scene, resulting in physical inconsistencies such as contact slippage or mesh penetration during terrain-aware tasks. In this work, we present MeshMimic, an innovative framework that bridges 3D scene reconstruction and embodied intelligence to enable humanoid robots to learn coupled "motion-terrain" interactions directly from video. By leveraging state-of-the-art 3D vision models, our framework precisely segments and reconstructs both human trajectories and the underlying 3D geometry of terrains and objects. We introduce an optimization algorithm based on kinematic consistency to extract high-quality motion data from noisy visual reconstructions, alongside a contact-invariant retargeting method that transfers human-environment interaction features to the humanoid agent. Experimental results demonstrate that MeshMimic achieves robust, highly dynamic performance across diverse and challenging terrains. Our approach proves that a low-cost pipeline utilizing only consumer-grade monocular sensors can facilitate the training of complex physical interactions, offering a scalable path toward the autonomous evolution of humanoid robots in unstructured environments.
format Preprint
id arxiv_https___arxiv_org_abs_2602_15733
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction
Zhang, Qiang
Ma, Jiahao
Liu, Peiran
Shi, Shuai
Su, Zeran
Wang, Zifan
Sun, Jingkai
Cui, Wei
Yu, Jialin
Han, Gang
Zhao, Wen
Sun, Pihai
Yin, Kangning
Wang, Jiaxu
Cao, Jiahang
Zhang, Lingfeng
Cheng, Hao
Hao, Xiaoshuai
Ji, Yiding
Liang, Junwei
Tang, Jian
Xu, Renjing
Guo, Yijie
Robotics
Artificial Intelligence
Humanoid motion control has witnessed significant breakthroughs in recent years, with deep reinforcement learning (RL) emerging as a primary catalyst for achieving complex, human-like behaviors. However, the high dimensionality and intricate dynamics of humanoid robots make manual motion design impractical, leading to a heavy reliance on expensive motion capture (MoCap) data. These datasets are not only costly to acquire but also frequently lack the necessary geometric context of the surrounding physical environment. Consequently, existing motion synthesis frameworks often suffer from a decoupling of motion and scene, resulting in physical inconsistencies such as contact slippage or mesh penetration during terrain-aware tasks. In this work, we present MeshMimic, an innovative framework that bridges 3D scene reconstruction and embodied intelligence to enable humanoid robots to learn coupled "motion-terrain" interactions directly from video. By leveraging state-of-the-art 3D vision models, our framework precisely segments and reconstructs both human trajectories and the underlying 3D geometry of terrains and objects. We introduce an optimization algorithm based on kinematic consistency to extract high-quality motion data from noisy visual reconstructions, alongside a contact-invariant retargeting method that transfers human-environment interaction features to the humanoid agent. Experimental results demonstrate that MeshMimic achieves robust, highly dynamic performance across diverse and challenging terrains. Our approach proves that a low-cost pipeline utilizing only consumer-grade monocular sensors can facilitate the training of complex physical interactions, offering a scalable path toward the autonomous evolution of humanoid robots in unstructured environments.
title MeshMimic: Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2602.15733