Scenario Understanding of Traffic Scenes Through Large Visual Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rivera, Esteban, Lübberstedt, Jannik, Uhlemann, Nico, Lienkamp, Markus
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917978598014976
author Rivera, Esteban
Lübberstedt, Jannik
Uhlemann, Nico
Lienkamp, Markus
author_facet Rivera, Esteban
Lübberstedt, Jannik
Uhlemann, Nico
Lienkamp, Markus
contents Deep learning models for autonomous driving, encompassing perception, planning, and control, depend on vast datasets to achieve their high performance. However, their generalization often suffers due to domain-specific data distributions, making an effective scene-based categorization of samples necessary to improve their reliability across diverse domains. Manual captioning, though valuable, is both labor-intensive and time-consuming, creating a bottleneck in the data annotation process. Large Visual Language Models (LVLMs) present a compelling solution by automating image analysis and categorization through contextual queries, often without requiring retraining for new categories. In this study, we evaluate the capabilities of LVLMs, including GPT-4 and LLaVA, to understand and classify urban traffic scenes on both an in-house dataset and the BDD100K. We propose a scalable captioning pipeline that integrates state-of-the-art models, enabling a flexible deployment on new datasets. Our analysis, combining quantitative metrics with qualitative insights, demonstrates the effectiveness of LVLMs to understand urban traffic scenarios and highlights their potential as an efficient tool for data-driven advancements in autonomous driving.
format Preprint
id arxiv_https___arxiv_org_abs_2501_17131
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Scenario Understanding of Traffic Scenes Through Large Visual Language Models
Rivera, Esteban
Lübberstedt, Jannik
Uhlemann, Nico
Lienkamp, Markus
Computer Vision and Pattern Recognition
Deep learning models for autonomous driving, encompassing perception, planning, and control, depend on vast datasets to achieve their high performance. However, their generalization often suffers due to domain-specific data distributions, making an effective scene-based categorization of samples necessary to improve their reliability across diverse domains. Manual captioning, though valuable, is both labor-intensive and time-consuming, creating a bottleneck in the data annotation process. Large Visual Language Models (LVLMs) present a compelling solution by automating image analysis and categorization through contextual queries, often without requiring retraining for new categories. In this study, we evaluate the capabilities of LVLMs, including GPT-4 and LLaVA, to understand and classify urban traffic scenes on both an in-house dataset and the BDD100K. We propose a scalable captioning pipeline that integrates state-of-the-art models, enabling a flexible deployment on new datasets. Our analysis, combining quantitative metrics with qualitative insights, demonstrates the effectiveness of LVLMs to understand urban traffic scenarios and highlights their potential as an efficient tool for data-driven advancements in autonomous driving.
title Scenario Understanding of Traffic Scenes Through Large Visual Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2501.17131