SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Rani, Anku, Dalal, Dwip, Gautam, Shreya, Gupta, Pankaj, Jain, Vinija, Chadha, Aman, Sheth, Amit, Das, Amitava
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918085367169024
author Rani, Anku
Dalal, Dwip
Gautam, Shreya
Gupta, Pankaj
Jain, Vinija
Chadha, Aman
Sheth, Amit
Das, Amitava
author_facet Rani, Anku
Dalal, Dwip
Gautam, Shreya
Gupta, Pankaj
Jain, Vinija
Chadha, Aman
Sheth, Amit
Das, Amitava
contents Deception is the intentional practice of twisting information. It is a nuanced societal practice deeply intertwined with human societal evolution, characterized by a multitude of facets. This research explores the problem of deception through the lens of psychology, employing a framework that categorizes deception into three forms: lies of omission, lies of commission, and lies of influence. The primary focus of this study is specifically on investigating only lies of omission. We propose a novel framework for deception detection leveraging NLP techniques. We curated an annotated dataset of 876,784 samples by amalgamating a popular large-scale fake news dataset and scraped news headlines from the Twitter handle of the Times of India, a well-known Indian news media house. Each sample has been labeled with four layers, namely: (i) the type of omission (speculation, bias, distortion, sounds factual, and opinion), (ii) colors of lies(black, white, etc), and (iii) the intention of such lies (to influence, etc) (iv) topic of lies (political, educational, religious, etc). We present a novel multi-task learning pipeline that leverages the dataless merging of fine-tuned language models to address the deception detection task mentioned earlier. Our proposed model achieved an F1 score of 0.87, demonstrating strong performance across all layers, including the type, color, intent, and topic aspects of deceptive content. Finally, our research explores the relationship between lies of omission and propaganda techniques. To accomplish this, we conducted an in-depth analysis, uncovering compelling findings. For instance, our analysis revealed a significant correlation between loaded language and opinion, shedding light on their interconnectedness. To encourage further research in this field, we are releasing the SEPSIS dataset and code at https://huggingface.co/datasets/ankurani/deception.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00292
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
Rani, Anku
Dalal, Dwip
Gautam, Shreya
Gupta, Pankaj
Jain, Vinija
Chadha, Aman
Sheth, Amit
Das, Amitava
Computation and Language
Deception is the intentional practice of twisting information. It is a nuanced societal practice deeply intertwined with human societal evolution, characterized by a multitude of facets. This research explores the problem of deception through the lens of psychology, employing a framework that categorizes deception into three forms: lies of omission, lies of commission, and lies of influence. The primary focus of this study is specifically on investigating only lies of omission. We propose a novel framework for deception detection leveraging NLP techniques. We curated an annotated dataset of 876,784 samples by amalgamating a popular large-scale fake news dataset and scraped news headlines from the Twitter handle of the Times of India, a well-known Indian news media house. Each sample has been labeled with four layers, namely: (i) the type of omission (speculation, bias, distortion, sounds factual, and opinion), (ii) colors of lies(black, white, etc), and (iii) the intention of such lies (to influence, etc) (iv) topic of lies (political, educational, religious, etc). We present a novel multi-task learning pipeline that leverages the dataless merging of fine-tuned language models to address the deception detection task mentioned earlier. Our proposed model achieved an F1 score of 0.87, demonstrating strong performance across all layers, including the type, color, intent, and topic aspects of deceptive content. Finally, our research explores the relationship between lies of omission and propaganda techniques. To accomplish this, we conducted an in-depth analysis, uncovering compelling findings. For instance, our analysis revealed a significant correlation between loaded language and opinion, shedding light on their interconnectedness. To encourage further research in this field, we are releasing the SEPSIS dataset and code at https://huggingface.co/datasets/ankurani/deception.
title SEPSIS: I Can Catch Your Lies -- A New Paradigm for Deception Detection
topic Computation and Language
url https://arxiv.org/abs/2312.00292