DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ishaq, Ayesha, Lahoud, Jean, More, Ketan, Thawakar, Omkar, Thawkar, Ritesh, Dissanayake, Dinura, Ahsan, Noor, Li, Yuhao, Khan, Fahad Shahbaz, Cholakkal, Hisham, Laptev, Ivan, Anwer, Rao Muhammad, Khan, Salman
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913734910279680
author Ishaq, Ayesha
Lahoud, Jean
More, Ketan
Thawakar, Omkar
Thawkar, Ritesh
Dissanayake, Dinura
Ahsan, Noor
Li, Yuhao
Khan, Fahad Shahbaz
Cholakkal, Hisham
Laptev, Ivan
Anwer, Rao Muhammad
Khan, Salman
author_facet Ishaq, Ayesha
Lahoud, Jean
More, Ketan
Thawakar, Omkar
Thawkar, Ritesh
Dissanayake, Dinura
Ahsan, Noor
Li, Yuhao
Khan, Fahad Shahbaz
Cholakkal, Hisham
Laptev, Ivan
Anwer, Rao Muhammad
Khan, Salman
contents While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive processing before decisions can be made. In this domain, a sequential and interpretive understanding of visual cues is essential for effective perception, prediction, and planning. Nevertheless, common VQA benchmarks often focus on the accuracy of the final answer while overlooking the reasoning process that enables the generation of accurate responses. Moreover, existing methods lack a comprehensive framework for evaluating step-by-step reasoning in realistic driving scenarios. To address this gap, we propose DriveLMM-o1, a new dataset and benchmark specifically designed to advance step-wise visual reasoning for autonomous driving. Our benchmark features over 18k VQA examples in the training set and more than 4k in the test set, covering diverse questions on perception, prediction, and planning, each enriched with step-by-step reasoning to ensure logical inference in autonomous driving scenarios. We further introduce a large multimodal model that is fine-tuned on our reasoning dataset, demonstrating robust performance in complex driving scenarios. In addition, we benchmark various open-source and closed-source methods on our proposed dataset, systematically comparing their reasoning capabilities for autonomous driving tasks. Our model achieves a +7.49% gain in final answer accuracy, along with a 3.62% improvement in reasoning score over the previous best open-source model. Our framework, dataset, and model are available at https://github.com/ayesha-ishaq/DriveLMM-o1.
format Preprint
id arxiv_https___arxiv_org_abs_2503_10621
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
Ishaq, Ayesha
Lahoud, Jean
More, Ketan
Thawakar, Omkar
Thawkar, Ritesh
Dissanayake, Dinura
Ahsan, Noor
Li, Yuhao
Khan, Fahad Shahbaz
Cholakkal, Hisham
Laptev, Ivan
Anwer, Rao Muhammad
Khan, Salman
Computer Vision and Pattern Recognition
Robotics
While large multimodal models (LMMs) have demonstrated strong performance across various Visual Question Answering (VQA) tasks, certain challenges require complex multi-step reasoning to reach accurate answers. One particularly challenging task is autonomous driving, which demands thorough cognitive processing before decisions can be made. In this domain, a sequential and interpretive understanding of visual cues is essential for effective perception, prediction, and planning. Nevertheless, common VQA benchmarks often focus on the accuracy of the final answer while overlooking the reasoning process that enables the generation of accurate responses. Moreover, existing methods lack a comprehensive framework for evaluating step-by-step reasoning in realistic driving scenarios. To address this gap, we propose DriveLMM-o1, a new dataset and benchmark specifically designed to advance step-wise visual reasoning for autonomous driving. Our benchmark features over 18k VQA examples in the training set and more than 4k in the test set, covering diverse questions on perception, prediction, and planning, each enriched with step-by-step reasoning to ensure logical inference in autonomous driving scenarios. We further introduce a large multimodal model that is fine-tuned on our reasoning dataset, demonstrating robust performance in complex driving scenarios. In addition, we benchmark various open-source and closed-source methods on our proposed dataset, systematically comparing their reasoning capabilities for autonomous driving tasks. Our model achieves a +7.49% gain in final answer accuracy, along with a 3.62% improvement in reasoning score over the previous best open-source model. Our framework, dataset, and model are available at https://github.com/ayesha-ishaq/DriveLMM-o1.
title DriveLMM-o1: A Step-by-Step Reasoning Dataset and Large Multimodal Model for Driving Scenario Understanding
topic Computer Vision and Pattern Recognition
Robotics
url https://arxiv.org/abs/2503.10621