Understanding Real-World Traffic Safety through RoadSafe365 Benchmark

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Xinyu, Jacob, Darryl C., Liu, Yuxin, Du, Xinsong, Ye, Muchao, Zhou, Bolei, He, Pan
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917255085817856
author Liu, Xinyu
Jacob, Darryl C.
Liu, Yuxin
Du, Xinsong
Ye, Muchao
Zhou, Bolei
He, Pan
author_facet Liu, Xinyu
Jacob, Darryl C.
Liu, Yuxin
Du, Xinsong
Ye, Muchao
Zhou, Bolei
He, Pan
contents Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale vision-language benchmark that supports fine-grained analysis of traffic safety from extensive and diverse real-world video data collections. Unlike prior works that focus primarily on coarse accident identification, RoadSafe365 is independently curated and systematically organized using a hierarchical taxonomy that refines and extends foundational definitions of crash, incident, and violation to bridge official traffic safety standards with data-driven traffic understanding systems. RoadSafe365 provides rich attribute annotations across diverse traffic event types, environmental contexts, and interaction scenarios, yielding 36,196 annotated clips from both dashcam and surveillance cameras. Each clip is paired with multiple-choice question-answer sets, comprising 864K candidate options, 8.4K unique answers, and 36K detailed scene descriptions collectively designed for vision-language understanding and reasoning. We establish strong baselines and observe consistent gains when fine-tuning on RoadSafe365. Cross-domain experiments on both real and synthetic datasets further validate its effectiveness. Designed for large-scale training and standardized evaluation, RoadSafe365 provides a comprehensive benchmark to advance reproducible research in real-world traffic safety analysis.
format Preprint
id arxiv_https___arxiv_org_abs_2602_07212
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Understanding Real-World Traffic Safety through RoadSafe365 Benchmark
Liu, Xinyu
Jacob, Darryl C.
Liu, Yuxin
Du, Xinsong
Ye, Muchao
Zhou, Bolei
He, Pan
Computer Vision and Pattern Recognition
Although recent traffic benchmarks have advanced multimodal data analysis, they generally lack systematic evaluation aligned with official safety standards. To fill this gap, we introduce RoadSafe365, a large-scale vision-language benchmark that supports fine-grained analysis of traffic safety from extensive and diverse real-world video data collections. Unlike prior works that focus primarily on coarse accident identification, RoadSafe365 is independently curated and systematically organized using a hierarchical taxonomy that refines and extends foundational definitions of crash, incident, and violation to bridge official traffic safety standards with data-driven traffic understanding systems. RoadSafe365 provides rich attribute annotations across diverse traffic event types, environmental contexts, and interaction scenarios, yielding 36,196 annotated clips from both dashcam and surveillance cameras. Each clip is paired with multiple-choice question-answer sets, comprising 864K candidate options, 8.4K unique answers, and 36K detailed scene descriptions collectively designed for vision-language understanding and reasoning. We establish strong baselines and observe consistent gains when fine-tuning on RoadSafe365. Cross-domain experiments on both real and synthetic datasets further validate its effectiveness. Designed for large-scale training and standardized evaluation, RoadSafe365 provides a comprehensive benchmark to advance reproducible research in real-world traffic safety analysis.
title Understanding Real-World Traffic Safety through RoadSafe365 Benchmark
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2602.07212