Whole-Body Conditioned Egocentric Video Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bai, Yutong, Tran, Danny, Bar, Amir, LeCun, Yann, Darrell, Trevor, Malik, Jitendra
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916811764662272
author Bai, Yutong
Tran, Danny
Bar, Amir
LeCun, Yann
Darrell, Trevor
Malik, Jitendra
author_facet Bai, Yutong
Tran, Danny
Bar, Amir
LeCun, Yann
Darrell, Trevor
Malik, Jitendra
contents We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions shape the environment from a first-person point of view. We train an auto-regressive conditional diffusion transformer on Nymeria, a large-scale dataset of real-world egocentric video and body pose capture. We further design a hierarchical evaluation protocol with increasingly challenging tasks, enabling a comprehensive analysis of the model's embodied prediction and control abilities. Our work represents an initial attempt to tackle the challenges of modeling complex real-world environments and embodied agent behaviors with video prediction from the perspective of a human.
format Preprint
id arxiv_https___arxiv_org_abs_2506_21552
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Whole-Body Conditioned Egocentric Video Prediction
Bai, Yutong
Tran, Danny
Bar, Amir
LeCun, Yann
Darrell, Trevor
Malik, Jitendra
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Robotics
We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions shape the environment from a first-person point of view. We train an auto-regressive conditional diffusion transformer on Nymeria, a large-scale dataset of real-world egocentric video and body pose capture. We further design a hierarchical evaluation protocol with increasingly challenging tasks, enabling a comprehensive analysis of the model's embodied prediction and control abilities. Our work represents an initial attempt to tackle the challenges of modeling complex real-world environments and embodied agent behaviors with video prediction from the perspective of a human.
title Whole-Body Conditioned Egocentric Video Prediction
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Multimedia
Robotics
url https://arxiv.org/abs/2506.21552