CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhou, Xingcheng, Guo, Hao, Song, Rui, Zimmer, Walter, Liu, Mingyu, Schamschurko, André, Cao, Hu, Knoll, Alois
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866910157382877184
author Zhou, Xingcheng
Guo, Hao
Song, Rui
Zimmer, Walter
Liu, Mingyu
Schamschurko, André
Cao, Hu
Knoll, Alois
author_facet Zhou, Xingcheng
Guo, Hao
Song, Rui
Zimmer, Walter
Liu, Mingyu
Schamschurko, André
Cao, Hu
Knoll, Alois
contents Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses under near-identical counterfactual scenes. We present CCTVBench, a Contrastive Consistency Traffic VideoQA Benchmark built on paired real accident videos and world-model-generated counterfactual counterparts, together with minimally different, mutually exclusive hypothesis questions. CCTVBench enforces a single structured decision pattern over each video question quadruple and provides actionable diagnostics that decompose failures into positive omission, positive swap, negative hallucination, and mutual-exclusivity violation, while separating video versus question consistency. Experiments across open-source and proprietary video LLMs reveal a large and persistent gap between standard per-instance QA metrics and quadruple-level contrastive consistency, with unreliable none-of-the-above rejection as a key bottleneck. Finally, we introduce C-TCD, a contrastive decoding approach leveraging a semantically exclusive counterpart video as the contrast input at inference time, improving both instance-level QA and contrastive consistency.
format Preprint
id arxiv_https___arxiv_org_abs_2604_20460
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Zhou, Xingcheng
Guo, Hao
Song, Rui
Zimmer, Walter
Liu, Mingyu
Schamschurko, André
Cao, Hu
Knoll, Alois
Computer Vision and Pattern Recognition
Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses under near-identical counterfactual scenes. We present CCTVBench, a Contrastive Consistency Traffic VideoQA Benchmark built on paired real accident videos and world-model-generated counterfactual counterparts, together with minimally different, mutually exclusive hypothesis questions. CCTVBench enforces a single structured decision pattern over each video question quadruple and provides actionable diagnostics that decompose failures into positive omission, positive swap, negative hallucination, and mutual-exclusivity violation, while separating video versus question consistency. Experiments across open-source and proprietary video LLMs reveal a large and persistent gap between standard per-instance QA metrics and quadruple-level contrastive consistency, with unreliable none-of-the-above rejection as a key bottleneck. Finally, we introduce C-TCD, a contrastive decoding approach leveraging a semantically exclusive counterpart video as the contrast input at inference time, improving both instance-level QA and contrastive consistency.
title CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2604.20460