MegaFake: A Theory-Driven Dataset of Fake News Generated by Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Lionel Z., Ng, Ka Chung, Ma, Yiming, Fan, Wenqi
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918439158808576
author Wang, Lionel Z.
Ng, Ka Chung
Ma, Yiming
Fan, Wenqi
author_facet Wang, Lionel Z.
Ng, Ka Chung
Ma, Yiming
Fan, Wenqi
contents Fake news significantly influences decision-making processes by misleading individuals, organizations, and even governments. Large language models (LLMs), as part of generative AI, can amplify this problem by generating highly convincing fake news at scale, posing a significant threat to online information integrity. Therefore, understanding the motivations and mechanisms behind fake news generated by LLMs is crucial for effective detection and governance. In this study, we develop the LLM-Fake Theory, a theoretical framework that integrates various social psychology theories to explain machine-generated deception. Guided by this framework, we design an innovative prompt engineering pipeline that automates fake news generation using LLMs, eliminating manual annotation needs. Utilizing this pipeline, we create a theoretically informed \underline{M}achin\underline{e}-\underline{g}ener\underline{a}ted \underline{Fake} news dataset, MegaFake, derived from FakeNewsNet. Through extensive experiments with MegaFake, we advance both theoretical understanding of human-machine deception mechanisms and practical approaches to fake news detection in the LLM era.
format Preprint
id arxiv_https___arxiv_org_abs_2408_11871
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle MegaFake: A Theory-Driven Dataset of Fake News Generated by Large Language Models
Wang, Lionel Z.
Ng, Ka Chung
Ma, Yiming
Fan, Wenqi
Computation and Language
Artificial Intelligence
Fake news significantly influences decision-making processes by misleading individuals, organizations, and even governments. Large language models (LLMs), as part of generative AI, can amplify this problem by generating highly convincing fake news at scale, posing a significant threat to online information integrity. Therefore, understanding the motivations and mechanisms behind fake news generated by LLMs is crucial for effective detection and governance. In this study, we develop the LLM-Fake Theory, a theoretical framework that integrates various social psychology theories to explain machine-generated deception. Guided by this framework, we design an innovative prompt engineering pipeline that automates fake news generation using LLMs, eliminating manual annotation needs. Utilizing this pipeline, we create a theoretically informed \underline{M}achin\underline{e}-\underline{g}ener\underline{a}ted \underline{Fake} news dataset, MegaFake, derived from FakeNewsNet. Through extensive experiments with MegaFake, we advance both theoretical understanding of human-machine deception mechanisms and practical approaches to fake news detection in the LLM era.
title MegaFake: A Theory-Driven Dataset of Fake News Generated by Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2408.11871