Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Yokoi, Shingo, Sasaki, Kento, Yamaguchi, Yu
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909844207828992
author Yokoi, Shingo
Sasaki, Kento
Yamaguchi, Yu
author_facet Yokoi, Shingo
Sasaki, Kento
Yamaguchi, Yu
contents Recent advances in end-to-end (E2E) autonomous driving have been enabled by training on diverse large-scale driving datasets, yet autonomous driving models still struggle in out-of-distribution (OOD) scenarios. The COOOL benchmark targets this gap by encouraging hazard understanding beyond closed taxonomies, and the 2COOOL challenge extends it to generating human-interpretable incident reports. We present a hierarchical reasoning framework for incident report generation from dashcam videos that integrates frame-level captioning, incident frame detection, and fine-grained reasoning within vision-language models (VLMs). We further improve factual accuracy and readability through model ensembling and a Blind A/B Scoring selection protocol. On the official 2COOOL open leaderboard, our method ranks 2nd among 29 teams and achieves the best CIDEr-D score, producing accurate and coherent incident narratives. These results indicate that hierarchical reasoning with VLMs is a promising direction for accident analysis and for broader understanding of safety-critical traffic events. The implementation and code are available at https://github.com/riron1206/kaggle-2COOOL-2nd-Place-Solution.
format Preprint
id arxiv_https___arxiv_org_abs_2510_12190
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
Yokoi, Shingo
Sasaki, Kento
Yamaguchi, Yu
Computer Vision and Pattern Recognition
Recent advances in end-to-end (E2E) autonomous driving have been enabled by training on diverse large-scale driving datasets, yet autonomous driving models still struggle in out-of-distribution (OOD) scenarios. The COOOL benchmark targets this gap by encouraging hazard understanding beyond closed taxonomies, and the 2COOOL challenge extends it to generating human-interpretable incident reports. We present a hierarchical reasoning framework for incident report generation from dashcam videos that integrates frame-level captioning, incident frame detection, and fine-grained reasoning within vision-language models (VLMs). We further improve factual accuracy and readability through model ensembling and a Blind A/B Scoring selection protocol. On the official 2COOOL open leaderboard, our method ranks 2nd among 29 teams and achieves the best CIDEr-D score, producing accurate and coherent incident narratives. These results indicate that hierarchical reasoning with VLMs is a promising direction for accident analysis and for broader understanding of safety-critical traffic events. The implementation and code are available at https://github.com/riron1206/kaggle-2COOOL-2nd-Place-Solution.
title Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2510.12190