Mitty: Diffusion-based Human-to-Robot Video Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Song, Yiren, Liu, Cheng, Mao, Weijia, Shou, Mike Zheng
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909970751029248
author Song, Yiren
Liu, Cheng
Mao, Weijia
Shou, Mike Zheng
author_facet Song, Yiren
Liu, Cheng
Mao, Weijia
Shou, Mike Zheng
contents Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss and cumulative errors that harm temporal and visual consistency. We present Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation. Built on a pretrained video diffusion model, Mitty leverages strong visual-temporal priors to translate human demonstrations into robot-execution videos without action labels or intermediate abstractions. Demonstration videos are compressed into condition tokens and fused with robot denoising tokens through bidirectional attention during diffusion. To mitigate paired-data scarcity, we also develop an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets. Experiments on Human2Robot and EPIC-Kitchens show that Mitty delivers state-of-the-art results, strong generalization to unseen environments, and new insights for scalable robot learning from human observations.
format Preprint
id arxiv_https___arxiv_org_abs_2512_17253
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Mitty: Diffusion-based Human-to-Robot Video Generation
Song, Yiren
Liu, Cheng
Mao, Weijia
Shou, Mike Zheng
Computer Vision and Pattern Recognition
Learning directly from human demonstration videos is a key milestone toward scalable and generalizable robot learning. Yet existing methods rely on intermediate representations such as keypoints or trajectories, introducing information loss and cumulative errors that harm temporal and visual consistency. We present Mitty, a Diffusion Transformer that enables video In-Context Learning for end-to-end Human2Robot video generation. Built on a pretrained video diffusion model, Mitty leverages strong visual-temporal priors to translate human demonstrations into robot-execution videos without action labels or intermediate abstractions. Demonstration videos are compressed into condition tokens and fused with robot denoising tokens through bidirectional attention during diffusion. To mitigate paired-data scarcity, we also develop an automatic synthesis pipeline that produces high-quality human-robot pairs from large egocentric datasets. Experiments on Human2Robot and EPIC-Kitchens show that Mitty delivers state-of-the-art results, strong generalization to unseen environments, and new insights for scalable robot learning from human observations.
title Mitty: Diffusion-based Human-to-Robot Video Generation
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.17253