Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sakajo, Haruki, Takato, Hiroshi, Tsutsui, Hiroshi, Soda, Komei, Kamigaito, Hidetaka, Watanabe, Taro
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915643958231040
author Sakajo, Haruki
Takato, Hiroshi
Tsutsui, Hiroshi
Soda, Komei
Kamigaito, Hidetaka
Watanabe, Taro
author_facet Sakajo, Haruki
Takato, Hiroshi
Tsutsui, Hiroshi
Soda, Komei
Kamigaito, Hidetaka
Watanabe, Taro
contents Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous driving. For example, LVLMs can generate safety-oriented descriptions of videos captured by road-facing cameras. However, ensuring comprehensive safety requires monitoring driver-facing views as well to detect risky events, such as the use of mobiles while driving. Thus, the ability to process synchronized inputs is necessary from both driver-facing and road-facing cameras. In this study, we develop models and investigate the capabilities of LVLMs by constructing a dataset and evaluating their performance on this dataset. Our experimental results demonstrate that while pre-trained LVLMs have limited effectiveness, fine-tuned LVLMs can generate accurate and safety-aware driving instructions. Nonetheless, several challenges remain, particularly in detecting subtle or complex events in the video. Our findings and error analysis provide valuable insights that can contribute to the improvement of LVLM-based systems in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2511_23311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
Sakajo, Haruki
Takato, Hiroshi
Tsutsui, Hiroshi
Soda, Komei
Kamigaito, Hidetaka
Watanabe, Taro
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Large-scale Vision Language Models (LVLMs) exhibit advanced capabilities in tasks that require visual information, including object detection. These capabilities have promising applications in various industrial domains, such as autonomous driving. For example, LVLMs can generate safety-oriented descriptions of videos captured by road-facing cameras. However, ensuring comprehensive safety requires monitoring driver-facing views as well to detect risky events, such as the use of mobiles while driving. Thus, the ability to process synchronized inputs is necessary from both driver-facing and road-facing cameras. In this study, we develop models and investigate the capabilities of LVLMs by constructing a dataset and evaluating their performance on this dataset. Our experimental results demonstrate that while pre-trained LVLMs have limited effectiveness, fine-tuned LVLMs can generate accurate and safety-aware driving instructions. Nonetheless, several challenges remain, particularly in detecting subtle or complex events in the video. Our findings and error analysis provide valuable insights that can contribute to the improvement of LVLM-based systems in this domain.
title Toward Automatic Safe Driving Instruction: A Large-Scale Vision Language Model Approach
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2511.23311