RED-DOT: Multimodal Fact-checking via Relevant Evidence Detection

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Papadopoulos, Stefanos-Iordanis, Koutlis, Christos, Papadopoulos, Symeon, Petrantonakis, Panagiotis C.
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866911790702526464
author Papadopoulos, Stefanos-Iordanis
Koutlis, Christos
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
author_facet Papadopoulos, Stefanos-Iordanis
Koutlis, Christos
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
contents Online misinformation is often multimodal in nature, i.e., it is caused by misleading associations between texts and accompanying images. To support the fact-checking process, researchers have been recently developing automatic multimodal methods that gather and analyze external information, evidence, related to the image-text pairs under examination. However, prior works assumed all external information collected from the web to be relevant. In this study, we introduce a "Relevant Evidence Detection" (RED) module to discern whether each piece of evidence is relevant, to support or refute the claim. Specifically, we develop the "Relevant Evidence Detection Directed Transformer" (RED-DOT) and explore multiple architectural variants (e.g., single or dual-stage) and mechanisms (e.g., "guided attention"). Extensive ablation and comparative experiments demonstrate that RED-DOT achieves significant improvements over the state-of-the-art (SotA) on the VERITE benchmark by up to 33.7%. Furthermore, our evidence re-ranking and element-wise modality fusion led to RED-DOT surpassing the SotA on NewsCLIPings+ by up to 3% without the need for numerous evidence or multiple backbone encoders. We release our code at: https://github.com/stevejpapad/relevant-evidence-detection
format Preprint
id arxiv_https___arxiv_org_abs_2311_09939
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle RED-DOT: Multimodal Fact-checking via Relevant Evidence Detection
Papadopoulos, Stefanos-Iordanis
Koutlis, Christos
Papadopoulos, Symeon
Petrantonakis, Panagiotis C.
Multimedia
Computer Vision and Pattern Recognition
Online misinformation is often multimodal in nature, i.e., it is caused by misleading associations between texts and accompanying images. To support the fact-checking process, researchers have been recently developing automatic multimodal methods that gather and analyze external information, evidence, related to the image-text pairs under examination. However, prior works assumed all external information collected from the web to be relevant. In this study, we introduce a "Relevant Evidence Detection" (RED) module to discern whether each piece of evidence is relevant, to support or refute the claim. Specifically, we develop the "Relevant Evidence Detection Directed Transformer" (RED-DOT) and explore multiple architectural variants (e.g., single or dual-stage) and mechanisms (e.g., "guided attention"). Extensive ablation and comparative experiments demonstrate that RED-DOT achieves significant improvements over the state-of-the-art (SotA) on the VERITE benchmark by up to 33.7%. Furthermore, our evidence re-ranking and element-wise modality fusion led to RED-DOT surpassing the SotA on NewsCLIPings+ by up to 3% without the need for numerous evidence or multiple backbone encoders. We release our code at: https://github.com/stevejpapad/relevant-evidence-detection
title RED-DOT: Multimodal Fact-checking via Relevant Evidence Detection
topic Multimedia
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.09939