HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Yinhuai, Zhao, Qihan, Lau, Yuen Fui, Yu, Runyi, Tsui, Hok Wai, Chen, Qifeng, Wang, Jingbo, Pang, Jiangmiao, Tan, Ping
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911416715313152
author Wang, Yinhuai
Zhao, Qihan
Lau, Yuen Fui
Yu, Runyi
Tsui, Hok Wai
Chen, Qifeng
Wang, Jingbo
Pang, Jiangmiao
Tan, Ping
author_facet Wang, Yinhuai
Zhao, Qihan
Lau, Yuen Fui
Yu, Runyi
Tsui, Hok Wai
Chen, Qifeng
Wang, Jingbo
Pang, Jiangmiao
Tan, Ping
contents Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need for meticulous, task-specific reward engineering, which limits their scalability. To narrow this gap, we present HumanX, a full-stack framework that compiles human video into generalizable, real-world interaction skills for humanoids, without task-specific rewards. HumanX integrates two co-designed components: XGen, a data generation pipeline that synthesizes diverse and physically plausible robot interaction data from video while supporting scalable data augmentation; and XMimic, a unified imitation learning framework that learns generalizable interaction skills. Evaluated across five distinct domains--basketball, football, badminton, cargo pickup, and reactive fighting--HumanX successfully acquires 10 different skills and transfers them zero-shot to a physical Unitree G1 humanoid. The learned capabilities include complex maneuvers such as pump-fake turnaround fadeaway jumpshots without any external perception, as well as interactive tasks like sustained human-robot passing sequences over 10 consecutive cycles--learned from a single video demonstration. Our experiments show that HumanX achieves over 8 times higher generalization success than prior methods, demonstrating a scalable and task-agnostic pathway for learning versatile, real-world robot interactive skills.
format Preprint
id arxiv_https___arxiv_org_abs_2602_02473
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos
Wang, Yinhuai
Zhao, Qihan
Lau, Yuen Fui
Yu, Runyi
Tsui, Hok Wai
Chen, Qifeng
Wang, Jingbo
Pang, Jiangmiao
Tan, Ping
Robotics
Machine Learning
Enabling humanoid robots to perform agile and adaptive interactive tasks has long been a core challenge in robotics. Current approaches are bottlenecked by either the scarcity of realistic interaction data or the need for meticulous, task-specific reward engineering, which limits their scalability. To narrow this gap, we present HumanX, a full-stack framework that compiles human video into generalizable, real-world interaction skills for humanoids, without task-specific rewards. HumanX integrates two co-designed components: XGen, a data generation pipeline that synthesizes diverse and physically plausible robot interaction data from video while supporting scalable data augmentation; and XMimic, a unified imitation learning framework that learns generalizable interaction skills. Evaluated across five distinct domains--basketball, football, badminton, cargo pickup, and reactive fighting--HumanX successfully acquires 10 different skills and transfers them zero-shot to a physical Unitree G1 humanoid. The learned capabilities include complex maneuvers such as pump-fake turnaround fadeaway jumpshots without any external perception, as well as interactive tasks like sustained human-robot passing sequences over 10 consecutive cycles--learned from a single video demonstration. Our experiments show that HumanX achieves over 8 times higher generalization success than prior methods, demonstrating a scalable and task-agnostic pathway for learning versatile, real-world robot interactive skills.
title HumanX: Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos
topic Robotics
Machine Learning
url https://arxiv.org/abs/2602.02473