GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866912158094196736 |
|---|---|
| author | Hassan, Mariam Stapf, Sebastian Rahimi, Ahmad Rezende, Pedro M B Haghighi, Yasaman Brüggemann, David Katircioglu, Isinsu Zhang, Lin Chen, Xiaoran Saha, Suman Cannici, Marco Aljalbout, Elie Ye, Botao Wang, Xi Davtyan, Aram Salzmann, Mathieu Scaramuzza, Davide Pollefeys, Marc Favaro, Paolo Alahi, Alexandre |
| author_facet | Hassan, Mariam Stapf, Sebastian Rahimi, Ahmad Rezende, Pedro M B Haghighi, Yasaman Brüggemann, David Katircioglu, Isinsu Zhang, Lin Chen, Xiaoran Saha, Suman Cannici, Marco Aljalbout, Elie Ye, Botao Wang, Xi Davtyan, Aram Salzmann, Mathieu Scaramuzza, Davide Pollefeys, Marc Favaro, Paolo Alahi, Alexandre |
| contents | We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_11198 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Hassan, Mariam Stapf, Sebastian Rahimi, Ahmad Rezende, Pedro M B Haghighi, Yasaman Brüggemann, David Katircioglu, Isinsu Zhang, Lin Chen, Xiaoran Saha, Suman Cannici, Marco Aljalbout, Elie Ye, Botao Wang, Xi Davtyan, Aram Salzmann, Mathieu Scaramuzza, Davide Pollefeys, Marc Favaro, Paolo Alahi, Alexandre Computer Vision and Pattern Recognition We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced. |
| title | GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.11198 |