CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2026
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866910157382877184 |
|---|---|
| author | Zhou, Xingcheng Guo, Hao Song, Rui Zimmer, Walter Liu, Mingyu Schamschurko, André Cao, Hu Knoll, Alois |
| author_facet | Zhou, Xingcheng Guo, Hao Song, Rui Zimmer, Walter Liu, Mingyu Schamschurko, André Cao, Hu Knoll, Alois |
| contents | Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses under near-identical counterfactual scenes. We present CCTVBench, a Contrastive Consistency Traffic VideoQA Benchmark built on paired real accident videos and world-model-generated counterfactual counterparts, together with minimally different, mutually exclusive hypothesis questions. CCTVBench enforces a single structured decision pattern over each video question quadruple and provides actionable diagnostics that decompose failures into positive omission, positive swap, negative hallucination, and mutual-exclusivity violation, while separating video versus question consistency. Experiments across open-source and proprietary video LLMs reveal a large and persistent gap between standard per-instance QA metrics and quadruple-level contrastive consistency, with unreliable none-of-the-above rejection as a key bottleneck. Finally, we introduce C-TCD, a contrastive decoding approach leveraging a semantically exclusive counterpart video as the contrast input at inference time, improving both instance-level QA and contrastive consistency. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_20460 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs Zhou, Xingcheng Guo, Hao Song, Rui Zimmer, Walter Liu, Mingyu Schamschurko, André Cao, Hu Knoll, Alois Computer Vision and Pattern Recognition Safety-critical traffic reasoning requires contrastive consistency: models must detect true hazards when an accident occurs, and reliably reject plausible-but-false hypotheses under near-identical counterfactual scenes. We present CCTVBench, a Contrastive Consistency Traffic VideoQA Benchmark built on paired real accident videos and world-model-generated counterfactual counterparts, together with minimally different, mutually exclusive hypothesis questions. CCTVBench enforces a single structured decision pattern over each video question quadruple and provides actionable diagnostics that decompose failures into positive omission, positive swap, negative hallucination, and mutual-exclusivity violation, while separating video versus question consistency. Experiments across open-source and proprietary video LLMs reveal a large and persistent gap between standard per-instance QA metrics and quadruple-level contrastive consistency, with unreliable none-of-the-above rejection as a key bottleneck. Finally, we introduce C-TCD, a contrastive decoding approach leveraging a semantically exclusive counterpart video as the contrast input at inference time, improving both instance-level QA and contrastive consistency. |
| title | CCTVBench: Contrastive Consistency Traffic VideoQA Benchmark for Multimodal LLMs |
| topic | Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2604.20460 |