Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929651109068800 |
|---|---|
| author | Picek, Lukas Čermák, Vojtěch Hanzl, Marek |
| author_facet | Picek, Lukas Čermák, Vojtěch Hanzl, Marek |
| contents | This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction detection, (ii) hazard object identification, and (iii) hazard captioning. We propose kernel-based change point detection on bounding boxes and optical flow dynamics for driver reaction detection to analyze motion patterns. For hazard identification, we combined a naive proximity-based strategy with object classification using a pre-trained ViT model. At last, for hazard captioning, we used the MOLMO vision-language model with tailored prompts to generate precise and context-aware descriptions of rare and low-resolution hazards. The proposed pipeline outperformed the baseline methods by a large margin, reducing the relative error by 33%, and scored 2nd on the final leaderboard consisting of 32 teams. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2412_19944 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark Picek, Lukas Čermák, Vojtěch Hanzl, Marek Computer Vision and Pattern Recognition This paper presents our submission to the COOOL competition, a novel benchmark for detecting and classifying out-of-label hazards in autonomous driving. Our approach integrates diverse methods across three core tasks: (i) driver reaction detection, (ii) hazard object identification, and (iii) hazard captioning. We propose kernel-based change point detection on bounding boxes and optical flow dynamics for driver reaction detection to analyze motion patterns. For hazard identification, we combined a naive proximity-based strategy with object classification using a pre-trained ViT model. At last, for hazard captioning, we used the MOLMO vision-language model with tailored prompts to generate precise and context-aware descriptions of rare and low-resolution hazards. The proposed pipeline outperformed the baseline methods by a large margin, reducing the relative error by 33%, and scored 2nd on the final leaderboard consisting of 32 teams. |
| title | Zero-shot Hazard Identification in Autonomous Driving: A Case Study on the COOOL Benchmark |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2412.19944 |