EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xu, Jilan, Huang, Yifei, Pei, Baoqi, Hou, Junlin, Li, Qingqiu, Chen, Guo, Zhang, Yuejie, Feng, Rui, Xie, Weidi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912330814586880
author Xu, Jilan
Huang, Yifei
Pei, Baoqi
Hou, Junlin
Li, Qingqiu
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
author_facet Xu, Jilan
Huang, Yifei
Pei, Baoqi
Hou, Junlin
Li, Qingqiu
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
contents Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate futur frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop an automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
format Preprint
id arxiv_https___arxiv_org_abs_2504_11732
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
Xu, Jilan
Huang, Yifei
Pei, Baoqi
Hou, Junlin
Li, Qingqiu
Chen, Guo
Zhang, Yuejie
Feng, Rui
Xie, Weidi
Computer Vision and Pattern Recognition
Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video, the first frame of the corresponding ego-centric video, and textual instructions, the goal is to generate futur frames of the ego-centric video. Inspired by the notion that hand-object interactions (HOI) in ego-centric videos represent the primary intentions and actions of the current actor, we present EgoExo-Gen that explicitly models the hand-object dynamics for cross-view video prediction. EgoExo-Gen consists of two stages. First, we design a cross-view HOI mask prediction model that anticipates the HOI masks in future ego-frames by modeling the spatio-temporal ego-exo correspondence. Next, we employ a video diffusion model to predict future ego-frames using the first ego-frame and textual instructions, while incorporating the HOI masks as structural guidance to enhance prediction quality. To facilitate training, we develop an automated pipeline to generate pseudo HOI masks for both ego- and exo-videos by exploiting vision foundation models. Extensive experiments demonstrate that our proposed EgoExo-Gen achieves better prediction performance compared to previous video prediction models on the Ego-Exo4D and H2O benchmark datasets, with the HOI masks significantly improving the generation of hands and interactive objects in the ego-centric videos.
title EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.11732