Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Trinci, Tomaso, Monteagudo, Henrique Piñeiro, Taccari, Leonardo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911704372215808
author Trinci, Tomaso
Monteagudo, Henrique Piñeiro
Taccari, Leonardo
author_facet Trinci, Tomaso
Monteagudo, Henrique Piñeiro
Taccari, Leonardo
contents Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to accurately perceive and reason about rare high-stakes dynamic events, such as collisions or near-collisions. To address this, we introduce a pipeline that enhances MLLM perception by fusing downsampled video frames with synchronized high-frequency telematics data (IMU and GPS) and semantic insights from specialized computer vision models. Our pipeline generates high-quality pseudo-labels, including descriptive captions and question-answer pairs, specifically designed to train MLLMs to identify and describe Safety-Critical Events (SCEs) in real-world driving footage. We show the effectiveness of our approach fine-tuning the open-source QwenVL-2.5 model via DoRA adapters: our experiments demonstrate significant improvements in identifying and explaining safety-critical events, with fewer than 50M trainable parameters and limited computational budget.
format Preprint
id arxiv_https___arxiv_org_abs_2605_22185
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis
Trinci, Tomaso
Monteagudo, Henrique Piñeiro
Taccari, Leonardo
Computer Vision and Pattern Recognition
Machine Learning
Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in general visual understanding. However, their application to safety-critical driving scenarios remains limited by an inability to accurately perceive and reason about rare high-stakes dynamic events, such as collisions or near-collisions. To address this, we introduce a pipeline that enhances MLLM perception by fusing downsampled video frames with synchronized high-frequency telematics data (IMU and GPS) and semantic insights from specialized computer vision models. Our pipeline generates high-quality pseudo-labels, including descriptive captions and question-answer pairs, specifically designed to train MLLMs to identify and describe Safety-Critical Events (SCEs) in real-world driving footage. We show the effectiveness of our approach fine-tuning the open-source QwenVL-2.5 model via DoRA adapters: our experiments demonstrate significant improvements in identifying and explaining safety-critical events, with fewer than 50M trainable parameters and limited computational budget.
title Enhancing Multimodal Large Language Models for Safety-Critical Driving Video Analysis
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2605.22185