Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Huang, Jen-tse, Chen, Chang, Lai, Shiyang, Wang, Wenxuan, Kaufman, Michelle R., Dredze, Mark
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916017151672320
author Huang, Jen-tse
Chen, Chang
Lai, Shiyang
Wang, Wenxuan
Kaufman, Michelle R.
Dredze, Mark
author_facet Huang, Jen-tse
Chen, Chang
Lai, Shiyang
Wang, Wenxuan
Kaufman, Michelle R.
Dredze, Mark
contents Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impressive reasoning capabilities, their robustness against misinformation entangled with cognitive biases remains under-explored. In this paper, we introduce a comprehensive evaluation framework using a high-quality, manually annotated dataset of 200 short videos spanning four health domains. This dataset provides fine-grained annotations for three deceptive patterns-experimental errors, logical fallacies, and fabricated claims-each verified by evidence such as national standards and academic literature. We evaluate eight frontier MLLMs across five modality settings. Experimental results demonstrate that Gemini-2.5-Pro achieves the highest performance in the multimodal setting with a belief score of 71.5/100, while o3 performs the worst at 35.2. Furthermore, we investigate social cues that induce false beliefs in videos and find that models are susceptible to biases like authoritative channel IDs.
format Preprint
id arxiv_https___arxiv_org_abs_2601_06600
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
Huang, Jen-tse
Chen, Chang
Lai, Shiyang
Wang, Wenxuan
Kaufman, Michelle R.
Dredze, Mark
Computation and Language
Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impressive reasoning capabilities, their robustness against misinformation entangled with cognitive biases remains under-explored. In this paper, we introduce a comprehensive evaluation framework using a high-quality, manually annotated dataset of 200 short videos spanning four health domains. This dataset provides fine-grained annotations for three deceptive patterns-experimental errors, logical fallacies, and fabricated claims-each verified by evidence such as national standards and academic literature. We evaluate eight frontier MLLMs across five modality settings. Experimental results demonstrate that Gemini-2.5-Pro achieves the highest performance in the multimodal setting with a belief score of 71.5/100, while o3 performs the worst at 35.2. Furthermore, we investigate social cues that induce false beliefs in videos and find that models are susceptible to biases like authoritative channel IDs.
title Probing Multimodal Large Language Models on Cognitive Biases in Chinese Short-Video Misinformation
topic Computation and Language
url https://arxiv.org/abs/2601.06600