RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sun, Zhenguo, Peng, Yibo, Meng, Yuan, Li, Xukun, Huang, Bo-Sheng, Bing, Zhenshan, Wang, Xinlong, Knoll, Alois
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912604156329984
author Sun, Zhenguo
Peng, Yibo
Meng, Yuan
Li, Xukun
Huang, Bo-Sheng
Bing, Zhenshan
Wang, Xinlong
Knoll, Alois
author_facet Sun, Zhenguo
Peng, Yibo
Meng, Yuan
Li, Xukun
Huang, Bo-Sheng
Bing, Zhenshan
Wang, Xinlong
Knoll, Alois
contents Long-horizon, high-dynamic motion tracking on humanoids remains brittle because absolute joint commands cannot compensate model-plant mismatch, leading to error accumulation. We propose RobotDancing, a simple, scalable framework that predicts residual joint targets to explicitly correct dynamics discrepancies. The pipeline is end-to-end--training, sim-to-sim validation, and zero-shot sim-to-real--and uses a single-stage reinforcement learning (RL) setup with a unified observation, reward, and hyperparameter configuration. We evaluate primarily on Unitree G1 with retargeted LAFAN1 dance sequences and validate transfer on H1/H1-2. RobotDancing can track multi-minute, high-energy behaviors (jumps, spins, cartwheels) and deploys zero-shot to hardware with high motion tracking quality.
format Preprint
id arxiv_https___arxiv_org_abs_2509_20717
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
Sun, Zhenguo
Peng, Yibo
Meng, Yuan
Li, Xukun
Huang, Bo-Sheng
Bing, Zhenshan
Wang, Xinlong
Knoll, Alois
Robotics
Artificial Intelligence
Long-horizon, high-dynamic motion tracking on humanoids remains brittle because absolute joint commands cannot compensate model-plant mismatch, leading to error accumulation. We propose RobotDancing, a simple, scalable framework that predicts residual joint targets to explicitly correct dynamics discrepancies. The pipeline is end-to-end--training, sim-to-sim validation, and zero-shot sim-to-real--and uses a single-stage reinforcement learning (RL) setup with a unified observation, reward, and hyperparameter configuration. We evaluate primarily on Unitree G1 with retargeted LAFAN1 dance sequences and validate transfer on H1/H1-2. RobotDancing can track multi-minute, high-energy behaviors (jumps, spins, cartwheels) and deploys zero-shot to hardware with high motion tracking quality.
title RobotDancing: Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking
topic Robotics
Artificial Intelligence
url https://arxiv.org/abs/2509.20717