MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bogdoll, Daniel, Yang, Yitian, Joseph, Tim, Yazgan, Melih, Zöllner, J. Marius
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908495728607232
author Bogdoll, Daniel
Yang, Yitian
Joseph, Tim
Yazgan, Melih
Zöllner, J. Marius
author_facet Bogdoll, Daniel
Yang, Yitian
Joseph, Tim
Yazgan, Melih
Zöllner, J. Marius
contents World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems. However, most works focus on camera data, with only a few that leverage lidar data or combine both to better represent autonomous vehicle sensor setups. In addition, raw sensor predictions are less actionable than 3D occupancy predictions, but there are no works examining the effects of combining both multimodal sensor data and 3D occupancy prediction. In this work, we perform a set of experiments with a MUltimodal World Model with Geometric VOxel representations (MUVO) to evaluate different sensor fusion strategies to better understand the effects on sensor data prediction. We also analyze potential weaknesses of current sensor fusion approaches and examine the benefits of additionally predicting 3D occupancy.
format Preprint
id arxiv_https___arxiv_org_abs_2311_11762
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations
Bogdoll, Daniel
Yang, Yitian
Joseph, Tim
Yazgan, Melih
Zöllner, J. Marius
Machine Learning
Robotics
World models for autonomous driving have the potential to dramatically improve the reasoning capabilities of today's systems. However, most works focus on camera data, with only a few that leverage lidar data or combine both to better represent autonomous vehicle sensor setups. In addition, raw sensor predictions are less actionable than 3D occupancy predictions, but there are no works examining the effects of combining both multimodal sensor data and 3D occupancy prediction. In this work, we perform a set of experiments with a MUltimodal World Model with Geometric VOxel representations (MUVO) to evaluate different sensor fusion strategies to better understand the effects on sensor data prediction. We also analyze potential weaknesses of current sensor fusion approaches and examine the benefits of additionally predicting 3D occupancy.
title MUVO: A Multimodal Generative World Model for Autonomous Driving with Geometric Representations
topic Machine Learning
Robotics
url https://arxiv.org/abs/2311.11762