OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xing, Shuo, Qian, Chengyuan, Wang, Yuping, Hua, Hongyuan, Tian, Kexin, Zhou, Yang, Tu, Zhengzhong
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916616650883072
author Xing, Shuo
Qian, Chengyuan
Wang, Yuping
Hua, Hongyuan
Tian, Kexin
Zhou, Yang
Tu, Zhengzhong
author_facet Xing, Shuo
Qian, Chengyuan
Wang, Yuping
Hua, Hongyuan
Tian, Kexin
Zhou, Yang
Tu, Zhengzhong
contents Since the advent of Multimodal Large Language Models (MLLMs), they have made a significant impact across a wide range of real-world applications, particularly in Autonomous Driving (AD). Their ability to process complex visual data and reason about intricate driving scenarios has paved the way for a new paradigm in end-to-end AD systems. However, the progress of developing end-to-end models for AD has been slow, as existing fine-tuning methods demand substantial resources, including extensive computational power, large-scale datasets, and significant funding. Drawing inspiration from recent advancements in inference computing, we propose OpenEMMA, an open-source end-to-end framework based on MLLMs. By incorporating the Chain-of-Thought reasoning process, OpenEMMA achieves significant improvements compared to the baseline when leveraging a diverse range of MLLMs. Furthermore, OpenEMMA demonstrates effectiveness, generalizability, and robustness across a variety of challenging driving scenarios, offering a more efficient and effective approach to autonomous driving. We release all the codes in https://github.com/taco-group/OpenEMMA.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15208
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
Xing, Shuo
Qian, Chengyuan
Wang, Yuping
Hua, Hongyuan
Tian, Kexin
Zhou, Yang
Tu, Zhengzhong
Computer Vision and Pattern Recognition
Machine Learning
Robotics
Since the advent of Multimodal Large Language Models (MLLMs), they have made a significant impact across a wide range of real-world applications, particularly in Autonomous Driving (AD). Their ability to process complex visual data and reason about intricate driving scenarios has paved the way for a new paradigm in end-to-end AD systems. However, the progress of developing end-to-end models for AD has been slow, as existing fine-tuning methods demand substantial resources, including extensive computational power, large-scale datasets, and significant funding. Drawing inspiration from recent advancements in inference computing, we propose OpenEMMA, an open-source end-to-end framework based on MLLMs. By incorporating the Chain-of-Thought reasoning process, OpenEMMA achieves significant improvements compared to the baseline when leveraging a diverse range of MLLMs. Furthermore, OpenEMMA demonstrates effectiveness, generalizability, and robustness across a variety of challenging driving scenarios, offering a more efficient and effective approach to autonomous driving. We release all the codes in https://github.com/taco-group/OpenEMMA.
title OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
topic Computer Vision and Pattern Recognition
Machine Learning
Robotics
url https://arxiv.org/abs/2412.15208