Saved in:
Bibliographic Details
Main Authors: Zhang, Yizhou, Trinh, Loc, Cao, Defu, Cui, Zijun, Liu, Yan
Format: Preprint
Published: 2023
Subjects:
Online Access:https://arxiv.org/abs/2304.07633
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911828695580672
author Zhang, Yizhou
Trinh, Loc
Cao, Defu
Cui, Zijun
Liu, Yan
author_facet Zhang, Yizhou
Trinh, Loc
Cao, Defu
Cui, Zijun
Liu, Yan
contents Recent years have witnessed the sustained evolution of misinformation that aims at manipulating public opinions. Unlike traditional rumors or fake news editors who mainly rely on generated and/or counterfeited images, text and videos, current misinformation creators now more tend to use out-of-context multimedia contents (e.g. mismatched images and captions) to deceive the public and fake news detection systems. This new type of misinformation increases the difficulty of not only detection but also clarification, because every individual modality is close enough to true information. To address this challenge, in this paper we explore how to achieve interpretable cross-modal de-contextualization detection that simultaneously identifies the mismatched pairs and the cross-modal contradictions, which is helpful for fact-check websites to document clarifications. The proposed model first symbolically disassembles the text-modality information to a set of fact queries based on the Abstract Meaning Representation of the caption and then forwards the query-image pairs into a pre-trained large vision-language model select the ``evidences" that are helpful for us to detect misinformation. Extensive experiments indicate that the proposed methodology can provide us with much more interpretable predictions while maintaining the accuracy same as the state-of-the-art model on this task.
format Preprint
id arxiv_https___arxiv_org_abs_2304_07633
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model
Zhang, Yizhou
Trinh, Loc
Cao, Defu
Cui, Zijun
Liu, Yan
Computation and Language
Machine Learning
Recent years have witnessed the sustained evolution of misinformation that aims at manipulating public opinions. Unlike traditional rumors or fake news editors who mainly rely on generated and/or counterfeited images, text and videos, current misinformation creators now more tend to use out-of-context multimedia contents (e.g. mismatched images and captions) to deceive the public and fake news detection systems. This new type of misinformation increases the difficulty of not only detection but also clarification, because every individual modality is close enough to true information. To address this challenge, in this paper we explore how to achieve interpretable cross-modal de-contextualization detection that simultaneously identifies the mismatched pairs and the cross-modal contradictions, which is helpful for fact-check websites to document clarifications. The proposed model first symbolically disassembles the text-modality information to a set of fact queries based on the Abstract Meaning Representation of the caption and then forwards the query-image pairs into a pre-trained large vision-language model select the ``evidences" that are helpful for us to detect misinformation. Extensive experiments indicate that the proposed methodology can provide us with much more interpretable predictions while maintaining the accuracy same as the state-of-the-art model on this task.
title Interpretable Detection of Out-of-Context Misinformation with Neural-Symbolic-Enhanced Large Multimodal Model
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2304.07633