ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Huang, Kung-Hsiang, Chan, Hou Pong, McKeown, Kathleen, Ji, Heng
Natura: Preprint
Pubblicazione: 2023
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912138154475520
author Huang, Kung-Hsiang
Chan, Hou Pong
McKeown, Kathleen
Ji, Heng
author_facet Huang, Kung-Hsiang
Chan, Hou Pong
McKeown, Kathleen
Ji, Heng
contents Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14225
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
Huang, Kung-Hsiang
Chan, Hou Pong
McKeown, Kathleen
Ji, Heng
Computation and Language
Considerable advancements have been made to tackle the misrepresentation of information derived from reference articles in the domains of fact-checking and faithful summarization. However, an unaddressed aspect remains - the identification of social media posts that manipulate information within associated news articles. This task presents a significant challenge, primarily due to the prevalence of personal opinions in such posts. We present a novel task, identifying manipulation of news on social media, which aims to detect manipulation in social media posts and identify manipulated or inserted information. To study this task, we have proposed a data collection schema and curated a dataset called ManiTweet, consisting of 3.6K pairs of tweets and corresponding articles. Our analysis demonstrates that this task is highly challenging, with large language models (LLMs) yielding unsatisfactory performance. Additionally, we have developed a simple yet effective basic model that outperforms LLMs significantly on the ManiTweet dataset. Finally, we have conducted an exploratory analysis of human-written tweets, unveiling intriguing connections between manipulation and the domain and factuality of news articles, as well as revealing that manipulated sentences are more likely to encapsulate the main story or consequences of a news outlet.
title ManiTweet: A New Benchmark for Identifying Manipulation of News on Social Media
topic Computation and Language
url https://arxiv.org/abs/2305.14225