Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Picek, Lukas, Čermák, Vojtěch, Hanzl, Marek
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929651109068800
author Picek, Lukas
Čermák, Vojtěch
Hanzl, Marek
author_facet Picek, Lukas
Čermák, Vojtěch
Hanzl, Marek
contents This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction detection, (ii) hazard object identification, and (iii) hazard captioning. We propose kernel-based change point detection on bounding boxes and optical flow dynamics for driver reaction detection to analyze motion patterns. For hazard identification, we combined a naive proximity-based strategy with object classification using a pre-trained ViT model. At last, for hazard captioning, we used the MOLMO vision-language model with tailored prompts to generate precise and context-aware descriptions of rare and low-resolution hazards. The proposed pipeline outperformed the baseline methods by a large margin, reducing the relative error by 33%, and scored 2nd on the final leaderboard consisting of 32 teams.
format Preprint
id arxiv_https___arxiv_org_abs_2412_19944
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark
Picek, Lukas
Čermák, Vojtěch
Hanzl, Marek
Computer Vision and Pattern Recognition
This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction detection, (ii) hazard object identification, and (iii) hazard captioning. We propose kernel-based change point detection on bounding boxes and optical flow dynamics for driver reaction detection to analyze motion patterns. For hazard identification, we combined a naive proximity-based strategy with object classification using a pre-trained ViT model. At last, for hazard captioning, we used the MOLMO vision-language model with tailored prompts to generate precise and context-aware descriptions of rare and low-resolution hazards. The proposed pipeline outperformed the baseline methods by a large margin, reducing the relative error by 33%, and scored 2nd on the final leaderboard consisting of 32 teams.
title Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2412.19944