Whole-Body Conditioned Egocentric Video Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916811764662272 |
|---|---|
| author | Bai, Yutong Tran, Danny Bar, Amir LeCun, Yann Darrell, Trevor Malik, Jitendra |
| author_facet | Bai, Yutong Tran, Danny Bar, Amir LeCun, Yann Darrell, Trevor Malik, Jitendra |
| contents | We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions shape the environment from a first-person point of view. We train an auto-regressive conditional diffusion transformer on Nymeria, a large-scale dataset of real-world egocentric video and body pose capture. We further design a hierarchical evaluation protocol with increasingly challenging tasks, enabling a comprehensive analysis of the model's embodied prediction and control abilities. Our work represents an initial attempt to tackle the challenges of modeling complex real-world environments and embodied agent behaviors with video prediction from the perspective of a human. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_21552 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Whole-Body Conditioned Egocentric Video Prediction Bai, Yutong Tran, Danny Bar, Amir LeCun, Yann Darrell, Trevor Malik, Jitendra Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Multimedia Robotics We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions shape the environment from a first-person point of view. We train an auto-regressive conditional diffusion transformer on Nymeria, a large-scale dataset of real-world egocentric video and body pose capture. We further design a hierarchical evaluation protocol with increasingly challenging tasks, enabling a comprehensive analysis of the model's embodied prediction and control abilities. Our work represents an initial attempt to tackle the challenges of modeling complex real-world environments and embodied agent behaviors with video prediction from the perspective of a human. |
| title | Whole-Body Conditioned Egocentric Video Prediction |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning Multimedia Robotics |
| url | https://arxiv.org/abs/2506.21552 |