OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yiren, Deng, Xiyao, Yang, Pei, Wang, Yihan, Shou, Mike Zheng
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916004901158912
author Song, Yiren
Deng, Xiyao
Yang, Pei
Wang, Yihan
Shou, Mike Zheng
author_facet Song, Yiren
Deng, Xiyao
Yang, Pei
Wang, Yihan
Shou, Mike Zheng
contents Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is that motion dynamics are partly transferable across embodiments, whereas appearance and morphology remain embodiment-specific. Existing approaches often entangle these factors, and many require paired data for every target embodiment, which limits scalability to new robots. We present OmniHumanoid, a framework that factorizes transferable motion learning and embodiment-specific adaptation. Our method learns a shared motion transfer model from motion-aligned paired videos spanning multiple embodiments, while adapting to a new embodiment using only unpaired videos through lightweight embodiment-specific adapters. To reduce interference between motion transfer and embodiment adaptation, we further introduce a branch-isolated attention design that separates motion conditioning from embodiment-specific modulation. In addition, we construct a synthetic cross-embodiment dataset with motion-aligned paired videos rendered across diverse humanoid assets, scenes, and viewpoints. Experiments on both synthetic and real-world benchmarks show that OmniHumanoid achieves strong motion fidelity and embodiment consistency, while enabling scalable adaptation to unseen humanoid embodiments without retraining the shared motion model.
format Preprint
id arxiv_https___arxiv_org_abs_2605_12038
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
Song, Yiren
Deng, Xiyao
Yang, Pei
Wang, Yihan
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Cross-embodiment video generation aims to transfer motions across different humanoid embodiments, such as human-to-robot and robot-to-robot, enabling scalable data generation for embodied intelligence. A major challenge in this setting is that motion dynamics are partly transferable across embodiments, whereas appearance and morphology remain embodiment-specific. Existing approaches often entangle these factors, and many require paired data for every target embodiment, which limits scalability to new robots. We present OmniHumanoid, a framework that factorizes transferable motion learning and embodiment-specific adaptation. Our method learns a shared motion transfer model from motion-aligned paired videos spanning multiple embodiments, while adapting to a new embodiment using only unpaired videos through lightweight embodiment-specific adapters. To reduce interference between motion transfer and embodiment adaptation, we further introduce a branch-isolated attention design that separates motion conditioning from embodiment-specific modulation. In addition, we construct a synthetic cross-embodiment dataset with motion-aligned paired videos rendered across diverse humanoid assets, scenes, and viewpoints. Experiments on both synthetic and real-world benchmarks show that OmniHumanoid achieves strong motion fidelity and embodiment consistency, while enabling scalable adaptation to unseen humanoid embodiments without retraining the shared motion model.
title OmniHumanoid: Streaming Cross-Embodiment Video Generation with Paired-Free Adaptation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.12038