Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zyblewski, Paweł, Klikowski, Jakub, Borek-Marciniec, Weronika, Ksieniewicz, Paweł
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918192476061696
author Zyblewski, Paweł
Klikowski, Jakub
Borek-Marciniec, Weronika
Ksieniewicz, Paweł
author_facet Zyblewski, Paweł
Klikowski, Jakub
Borek-Marciniec, Weronika
Ksieniewicz, Paweł
contents Tabular data is considered the last unconquered castle of deep learning, yet the task of data stream classification is stated to be an equally important and demanding research area. Due to the temporal constraints, it is assumed that deep learning methods are not the optimal solution for application in this field. However, excluding the entire -- and prevalent -- group of methods seems rather rash given the progress that has been made in recent years in its development. For this reason, the following paper is the first to present an approach to natural language data stream classification using the sentence space method, which allows for encoding text into the form of a discrete digital signal. This allows the use of convolutional deep networks dedicated to image classification to solve the task of recognizing fake news based on text data. Based on the real-life Fakeddit dataset, the proposed approach was compared with state-of-the-art algorithms for data stream classification based on generalization ability and time complexity.
format Preprint
id arxiv_https___arxiv_org_abs_2407_10807
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain
Zyblewski, Paweł
Klikowski, Jakub
Borek-Marciniec, Weronika
Ksieniewicz, Paweł
Computation and Language
Machine Learning
Tabular data is considered the last unconquered castle of deep learning, yet the task of data stream classification is stated to be an equally important and demanding research area. Due to the temporal constraints, it is assumed that deep learning methods are not the optimal solution for application in this field. However, excluding the entire -- and prevalent -- group of methods seems rather rash given the progress that has been made in recent years in its development. For this reason, the following paper is the first to present an approach to natural language data stream classification using the sentence space method, which allows for encoding text into the form of a discrete digital signal. This allows the use of convolutional deep networks dedicated to image classification to solve the task of recognizing fake news based on text data. Based on the real-life Fakeddit dataset, the proposed approach was compared with state-of-the-art algorithms for data stream classification based on generalization ability and time complexity.
title Employing Sentence Space Embedding for Classification of Data Stream from Fake News Domain
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2407.10807