CLASH: A Benchmark for Cross-Modal Contradiction Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Popordanoska, Teodora, Li, Jiameng, Blaschko, Matthew B.
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866917101365624832
author Popordanoska, Teodora
Li, Jiameng
Blaschko, Matthew B.
author_facet Popordanoska, Teodora
Li, Jiameng
Blaschko, Matthew B.
contents Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations and ensuring reliability. We introduce CLASH, a novel benchmark for multimodal contradiction detection, featuring COCO images paired with contradictory captions containing controlled object-level or attribute-level contradictions. The samples include targeted questions evaluated in both multiple-choice and open-ended formats. The benchmark provides an extensive fine-tuning set filtered through automated quality checks, alongside a smaller human-verified diagnostic set. Our analysis of state-of-the-art models reveals substantial limitations in recognizing cross-modal conflicts, exposing systematic modality biases and category-specific weaknesses. Furthermore, we empirically demonstrate that targeted fine-tuning on CLASH substantially enhances conflict detection capabilities.
format Preprint
id arxiv_https___arxiv_org_abs_2511_19199
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CLASH: A Benchmark for Cross-Modal Contradiction Detection
Popordanoska, Teodora
Li, Jiameng
Blaschko, Matthew B.
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Contradictory multimodal inputs are common in real-world settings, yet existing benchmarks typically assume input consistency and fail to evaluate cross-modal contradiction detection - a fundamental capability for preventing hallucinations and ensuring reliability. We introduce CLASH, a novel benchmark for multimodal contradiction detection, featuring COCO images paired with contradictory captions containing controlled object-level or attribute-level contradictions. The samples include targeted questions evaluated in both multiple-choice and open-ended formats. The benchmark provides an extensive fine-tuning set filtered through automated quality checks, alongside a smaller human-verified diagnostic set. Our analysis of state-of-the-art models reveals substantial limitations in recognizing cross-modal conflicts, exposing systematic modality biases and category-specific weaknesses. Furthermore, we empirically demonstrate that targeted fine-tuning on CLASH substantially enhances conflict detection capabilities.
title CLASH: A Benchmark for Cross-Modal Contradiction Detection
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2511.19199