GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Hassan, Mariam, Stapf, Sebastian, Rahimi, Ahmad, Rezende, Pedro M B, Haghighi, Yasaman, Brüggemann, David, Katircioglu, Isinsu, Zhang, Lin, Chen, Xiaoran, Saha, Suman, Cannici, Marco, Aljalbout, Elie, Ye, Botao, Wang, Xi, Davtyan, Aram, Salzmann, Mathieu, Scaramuzza, Davide, Pollefeys, Marc, Favaro, Paolo, Alahi, Alexandre
Format: Preprint
Publié: 2024
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912158094196736
author Hassan, Mariam
Stapf, Sebastian
Rahimi, Ahmad
Rezende, Pedro M B
Haghighi, Yasaman
Brüggemann, David
Katircioglu, Isinsu
Zhang, Lin
Chen, Xiaoran
Saha, Suman
Cannici, Marco
Aljalbout, Elie
Ye, Botao
Wang, Xi
Davtyan, Aram
Salzmann, Mathieu
Scaramuzza, Davide
Pollefeys, Marc
Favaro, Paolo
Alahi, Alexandre
author_facet Hassan, Mariam
Stapf, Sebastian
Rahimi, Ahmad
Rezende, Pedro M B
Haghighi, Yasaman
Brüggemann, David
Katircioglu, Isinsu
Zhang, Lin
Chen, Xiaoran
Saha, Suman
Cannici, Marco
Aljalbout, Elie
Ye, Botao
Wang, Xi
Davtyan, Aram
Salzmann, Mathieu
Scaramuzza, Davide
Pollefeys, Marc
Favaro, Paolo
Alahi, Alexandre
contents We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced.
format Preprint
id arxiv_https___arxiv_org_abs_2412_11198
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Hassan, Mariam
Stapf, Sebastian
Rahimi, Ahmad
Rezende, Pedro M B
Haghighi, Yasaman
Brüggemann, David
Katircioglu, Isinsu
Zhang, Lin
Chen, Xiaoran
Saha, Suman
Cannici, Marco
Aljalbout, Elie
Ye, Botao
Wang, Xi
Davtyan, Aram
Salzmann, Mathieu
Scaramuzza, Davide
Pollefeys, Marc
Favaro, Paolo
Alahi, Alexandre
Computer Vision and Pattern Recognition
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced.
title GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.11198