Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Masri, Sari, Ashqar, Huthaifa I., Elhenawy, Mohammed
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929735290847232
author Masri, Sari
Ashqar, Huthaifa I.
Elhenawy, Mohammed
author_facet Masri, Sari
Ashqar, Huthaifa I.
Elhenawy, Mohammed
contents Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such as GPT-4o, to provide logical and visual reasoning by directly using birds-eye-view videos of four-legged intersections. In this proposed method, GPT-4o acts as intelligent system to detect conflicts and provide explanations and recommendations for the drivers. The fine-tuned model achieved an accuracy of 77.14%, while the manual evaluation of the true predicted values of the fine-tuned GPT-4o showed significant achievements of 89.9% accuracy for model-generated explanations and 92.3% for the recommended next actions. These results highlight the feasibility of using MLLMs for real-time traffic management using videos as inputs, offering scalable and actionable insights into intersections traffic management and operation. Code used in this study is available at https://github.com/sarimasri3/Traffic-Intersection-Conflict-Detection-using-images.git.
format Preprint
id arxiv_https___arxiv_org_abs_2502_20573
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
Masri, Sari
Ashqar, Huthaifa I.
Elhenawy, Mohammed
Computer Vision and Pattern Recognition
Computation and Language
Traffic control in unsignalized urban intersections presents significant challenges due to the complexity, frequent conflicts, and blind spots. This study explores the capability of leveraging Multimodal Large Language Models (MLLMs), such as GPT-4o, to provide logical and visual reasoning by directly using birds-eye-view videos of four-legged intersections. In this proposed method, GPT-4o acts as intelligent system to detect conflicts and provide explanations and recommendations for the drivers. The fine-tuned model achieved an accuracy of 77.14%, while the manual evaluation of the true predicted values of the fine-tuned GPT-4o showed significant achievements of 89.9% accuracy for model-generated explanations and 92.3% for the recommended next actions. These results highlight the feasibility of using MLLMs for real-time traffic management using videos as inputs, offering scalable and actionable insights into intersections traffic management and operation. Code used in this study is available at https://github.com/sarimasri3/Traffic-Intersection-Conflict-Detection-using-images.git.
title Visual Reasoning at Urban Intersections: FineTuning GPT-4o for Traffic Conflict Detection
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2502.20573