Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Roux, Constant, De Matteïs, Ludovic, Jordana, Armand, Guillet, Valentin, Mansard, Nicolas, Stasse, Olivier, Souères, Philippe
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916039658307584
author Roux, Constant
De Matteïs, Ludovic
Jordana, Armand
Guillet, Valentin
Mansard, Nicolas
Stasse, Olivier
Souères, Philippe
author_facet Roux, Constant
De Matteïs, Ludovic
Jordana, Armand
Guillet, Valentin
Mansard, Nicolas
Stasse, Olivier
Souères, Philippe
contents Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches. Standard approaches rely on Geometric Retargeting or Indirect Dynamic Retargeting pipelines. We identify that these intermediate kinematic projections introduce a geometric bias, restricting the search space and yielding suboptimal dynamic behaviors. In this paper, we propose Direct Dynamic Retargeting (DDR), a novel single-stage framework that generates high-fidelity, dynamically feasible trajectories directly from expert videos. By formulating the problem in the task space and leveraging a sampling-based Model Predictive Control solver within a physics simulator, DDR natively optimizes over complex contact sequences while mitigating input drift. Our experiments demonstrate that bypassing the geometric bias allows DDR to outperform state-of-the-art baselines in demonstration tracking accuracy. Furthermore, we establish that providing such physically viable references to RL agents accelerates training convergence and enhances the final execution of agile and balancing behaviors. Source code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2605_23762
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos
Roux, Constant
De Matteïs, Ludovic
Jordana, Armand
Guillet, Valentin
Mansard, Nicolas
Stasse, Olivier
Souères, Philippe
Robotics
Imitation Learning from monocular video demonstrations provides a scalable approach for teaching complex skills to humanoid robots. However, translating human motion to humanoids requires overcoming significant morphological mismatches. Standard approaches rely on Geometric Retargeting or Indirect Dynamic Retargeting pipelines. We identify that these intermediate kinematic projections introduce a geometric bias, restricting the search space and yielding suboptimal dynamic behaviors. In this paper, we propose Direct Dynamic Retargeting (DDR), a novel single-stage framework that generates high-fidelity, dynamically feasible trajectories directly from expert videos. By formulating the problem in the task space and leveraging a sampling-based Model Predictive Control solver within a physics simulator, DDR natively optimizes over complex contact sequences while mitigating input drift. Our experiments demonstrate that bypassing the geometric bias allows DDR to outperform state-of-the-art baselines in demonstration tracking accuracy. Furthermore, we establish that providing such physically viable references to RL agents accelerates training convergence and enhances the final execution of agile and balancing behaviors. Source code will be made publicly available.
title Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos
topic Robotics
url https://arxiv.org/abs/2605.23762